llama.cpp
7c908502
- ggml-webgpu: improve MTP inference by using mat-vec path for small batches (#24811)
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
32 days ago
ggml-webgpu: improve MTP inference by using mat-vec path for small batches (#24811) * ggml-webgpu: improve small batches decoding * Add barrier to the NUM_COLS loop in mul-mat-vec
References
#24811 - ggml-webgpu: improve MTP inference by using mat-vec path for small batches
Author
yomaytk
Parents
035cd8f9
Loading