CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover #26079
CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover
ff05098b
Added Blackwell specific switch point, to reduce dependence on runtim…
f2a3efae
praneshgo
changed the title CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover 29 days ago
Add per-HW switch point values for DGX Spark and removing runtime env…
f2aace7f
praneshgo
force pushed
to
f2aace7f
28 days ago
Adding switch points for Ada, tested on RTX 4090
db1508d8
Modifying DGX Spark numbers based on latest run and adding some comme…
9052a5ce
praneshgo
marked this pull request as ready for review 24 days ago
ORippler
requested changes
on 2026-08-18
ORippler
requested changes
on 2026-08-19
Reverting an unnecessary conditional
614d5244
ORippler
approved these changes
on 2026-08-19
Update ggml/src/ggml-cuda/mmvq.cu
41a70d7e
ORippler
merged
2b562109
into master 7 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub