llama.cpp
CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (recreated)
#28552
Open

CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (recreated) #28552

pwilkin wants to merge 2 commits into master from pr-24546-recreated
pwilkin
ravel7524 CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 …
79800853
pwilkin CUDA: pick MMQ tile size against ncols_opt set on the host side
f01a7549
pwilkin pwilkin requested a review 1 day ago
JohannesGaessler
JohannesGaessler commented on 2026-09-07
JohannesGaessler JohannesGaessler assigned JohannesGaessler JohannesGaessler 1 day ago
github-actions github-actions added ggml
github-actions github-actions added CUDA
JohannesGaessler
JohannesGaessler
JohannesGaessler approved these changes on 2026-09-08

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
Labels
Milestone