llama.cpp
ggml-webgpu: improve MTP inference by using mat-vec path for small batches
#24811
Merged

ggml-webgpu: improve MTP inference by using mat-vec path for small batches #24811

yomaytk merged 2 commits into ggml-org:master from yomaytk:mat-vec-n8
yomaytk
yomaytk ggml-webgpu: improve small batches decoding
313ac645
yomaytk yomaytk requested a review from ggerganov ggerganov 41 days ago
yomaytk yomaytk requested a review 41 days ago
github-actions github-actions added testing
github-actions github-actions added ggml
github-actions github-actions added WebGPU
yomaytk
yomaytk commented on 2026-06-19
yomaytk
yomaytk commented on 2026-06-19
yomaytk
CISC
CISC approved these changes on 2026-06-19
CISC
yomaytk
yomaytk Add barrier to the NUM_COLS loop in mul-mat-vec
9fe606fe
yomaytk
reeselevine
reeselevine
reeselevine approved these changes on 2026-06-23
yomaytk
yomaytk yomaytk merged 7c908502 into master 37 days ago
yomaytk yomaytk deleted the mat-vec-n8 branch 37 days ago

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone