whisper.cpp
d9bf63cf - CUDA: add a fused top-K MoE kernel (llama/16130)

Commit

156 days ago

CUDA: add a fused top-K MoE kernel (llama/16130) * CUDA: add a fused top-K MoE kernel This kernel does the following: 1. softmax over the logits per token [n_experts, n_tokens] 2. argmax reduce over the top-k (n_experts_used) logits 3. write weights + ids to global memory It is intended as fusion of softmax->top-k->get_rows pipeline for MoE models * Refactor into ggml_cuda_should_use_topk_moe * Review: Use better coalescing pattern, use WARP_SIZE, store logits into registers before * Review: format + micro-optimizations * Fix bug: fix tie breakers * Add optional norm + clean-up code * Use smem for final write * Add bounds check * Use better memory pattern for writeback

References

#3428 - sync : ggml

Author

am17an

Committer

ggerganov

Parents

24ea5476

whisper.cpp d9bf63cf - CUDA: add a fused top-K MoE kernel (llama/16130)

whisper.cpp
d9bf63cf - CUDA: add a fused top-K MoE kernel (llama/16130)