llama.cpp
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion
#26947
Merged
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Overview
Commits
1
Changes
View On
GitHub
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion
#26947
ggerganov
merged 1 commit into
ggml-org:master
from
jinzihao:flash-attn-f16-f32
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion
f8503dd9
jinzihao
requested a review
from
ggerganov
11 days ago
ggml-gh-bot
added
draft
github-actions
added
ggml
github-actions
marked this pull request as draft
11 days ago
github-actions
removed
draft
jinzihao
marked this pull request as ready for review
11 days ago
am17an
approved these changes on 2026-08-12
am17an
added
merge ready
ggerganov
merged
eeae28b6
into master
10 days ago
Login to write a write a comment.
Login via GitHub
Reviewers
am17an
ggerganov
Assignees
No one assigned
Labels
ggml
merge ready
Milestone
No milestone
Login to write a write a comment.
Login via GitHub