llama.cpp
7c908502 - ggml-webgpu: improve MTP inference by using mat-vec path for small batches (#24811)

Commit
32 days ago
ggml-webgpu: improve MTP inference by using mat-vec path for small batches (#24811) * ggml-webgpu: improve small batches decoding * Add barrier to the NUM_COLS loop in mul-mat-vec
Author
Parents
Loading