llama.cpp
6e9007ae
- ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (#24530)
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
38 days ago
ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (#24530) * Improve prefill speeds for i-quants * Fix #if defined() usage in preprocessor guards.
References
#24530 - ggml-webgpu: improve i-quants mul_mat performance and speed up prefill
Author
yomaytk
Parents
dd4623a7
Loading