CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover (#26079)
* CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover
Add a runtime override of the mul_mat_vec_q -> MMQ batch crossover
(default MMVQ_MAX_BATCH_SIZE). Lowering it routes batches above the
threshold from the CUDA-core vector kernel to the int8 MMQ tensor-core
path, which is faster once quantized decode becomes compute-bound at
B>1 (measured +23-41% at B=8 on RTX 5090 for Q4_K dense, no low-batch loss).
The value is parsed once and clamped to [1, MMVQ_MAX_BATCH_SIZE], since
mul_mat_vec_q asserts ncols_dst <= that; invalid input warns and falls
back to the default. The override is applied consistently in both the
mul_mat_vec_q and MUL_MAT_ID dispatch paths. Default behavior unchanged.
* Added Blackwell specific switch point, to reduce dependence on runtime env var.
* Add per-HW switch point values for DGX Spark and removing runtime env var
* Adding switch points for Ada, tested on RTX 4090
* Modifying DGX Spark numbers based on latest run and adding some comments and small functional changes relating to MoE
* Reverting an unnecessary conditional
* Update ggml/src/ggml-cuda/mmvq.cu
---------
Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com>
Co-authored-by: Oliver Simons <osimons@nvidia.com>