llama.cpp
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion
#26947
Merged

ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion #26947

jinzihao
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion
f8503dd9
jinzihao jinzihao requested a review from ggerganov ggerganov 11 days ago
ggml-gh-bot
ggml-gh-bot ggml-gh-bot added draft
github-actions github-actions added ggml
github-actions github-actions marked this pull request as draft 11 days ago
github-actions github-actions removed draft
jinzihao jinzihao marked this pull request as ready for review 11 days ago
am17an
am17an approved these changes on 2026-08-12
am17an am17an added merge ready
ggerganov ggerganov merged eeae28b6 into master 10 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone