llama.cpp
6e9007ae - ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (#24530)

Commit
38 days ago
ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (#24530) * Improve prefill speeds for i-quants * Fix #if defined() usage in preprocessor guards.
Author
Parents
Loading