llama.cpp
cuda: unblock mmq for MoE on sm_60
#26264
Merged

cuda: unblock mmq for MoE on sm_60 #26264

am17an merged 3 commits into ggml-org:master from dfriehs:p100-moe-mmq
dfriehs
dfriehs dfriehs requested a review 34 days ago
github-actions github-actions added ggml
github-actions github-actions added CUDA
am17an
am17an approved these changes on 2026-07-29
mb8565
dfriehs
dfriehs dfriehs marked this pull request as draft 33 days ago
dfriehs
dfriehs dfriehs force pushed to af8a1ad5 32 days ago
dfriehs
dfriehs dfriehs marked this pull request as ready for review 32 days ago
dfriehs
am17an
Kenzu
dfriehs cuda: unblock mmq for MoE on sm_60
ddf75a2f
dfriehs cuda: duplicate mmq-config-pascal for dp4a and older
70f323f0
dfriehs cuda: reduce occupancy on non-dp4a pascal for Q2_K, Q4_K, Q5_K, Q6_K
1809286c
dfriehs dfriehs force pushed from af8a1ad5 to 1809286c 7 days ago
dfriehs
IMbackK
IMbackK approved these changes on 2026-08-26
am17an
am17an approved these changes on 2026-08-26
am17an am17an merged fc35562b into master 6 days ago
mb8565
mb8565

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone