llama.cpp
7fe2ae45 - sycl : port multi-column MMVQ from CUDA backend (#21845)

Commit

30 days ago

sycl : port multi-column MMVQ from CUDA backend (#21845) mmvq: Port the ncols_dst optimization from ggml-cuda/mmvq.cu to SYCL. Read weights once per dispatch instead of once per column. Covers all standard quant types + reorder paths for Q4_0, Q8_0, Q3_K, Q4_K, Q5_K, Q6_K. IQ types (except IQ4_XS) excluded due to incompatible vec_dot signatures. ggml-sycl: The weight reorder was only bootstrapped on single-token mat-vec (ne[1] == 1). Speculative / MTP verify issues only multi-column mat-vec, so it never triggered the reorder and ran on the slower non-reorder kernel. Bootstrap it on small multi-column batches (ne[1] <= 8) too.

References

#21845 - sycl : port multi-column MMVQ from CUDA backend (~45% speculative decoding speedup on Intel Arc)

Author

masonmilby

Parents

7c158fbb

llama.cpp 7fe2ae45 - sycl : port multi-column MMVQ from CUDA backend (#21845)

llama.cpp
7fe2ae45 - sycl : port multi-column MMVQ from CUDA backend (#21845)