llama.cpp
CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (recreated)
#28552
Merged

Commits
  • CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (#24546)
    pwilkin committed 4 days ago
  • CUDA: pick MMQ tile size against ncols_opt set on the host side
    pwilkin committed 4 days ago
Loading