llama.cpp
ggml-webgpu: FlashAttention refactor + standardize quantization support
#23834
Merged

ggml-webgpu: FlashAttention refactor + standardize quantization support #23834

reeselevine
reeselevine Start work on flash_attn refactor
e1a0bb34
reeselevine Refactor
8d61b5cd
reeselevine Split k/v quantization
19f150a5
reeselevine Refactor and abstract quantization logic for flash_attn and mul_mat
94a1fdde
reeselevine Add quantization support to tile path
ddedf2fc
reeselevine formatting
ffbd86f0
github-actions github-actions added ggml
github-actions github-actions added WebGPU
reeselevine Merge remote-tracking branch 'origin/master' into flash_attn_refactor
2e1d2b5c
reeselevine Merge remote-tracking branch 'upstream/master' into flash_attn_refactor
e0868e34
reeselevine reeselevine marked this pull request as ready for review 111 days ago
reeselevine reeselevine requested a review 111 days ago
reeselevine
yomaytk
yomaytk commented on 2026-06-02
yomaytk
yomaytk commented on 2026-06-02
yomaytk
yomaytk commented on 2026-06-02
reeselevine Move to functions, add a check
cf66f0a3
reeselevine reeselevine requested a review from ggerganov ggerganov 110 days ago
github-actions github-actions added testing
Constannnnnt
yomaytk
yomaytk approved these changes on 2026-06-03
reeselevine reeselevine added merge ready
ggerganov ggerganov merged e8c54893 into master 109 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone