llama.cpp
ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86
#29423
Open

ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 #29423

SongXiaoXi wants to merge 2 commits into ggml-org:master from SongXiaoXi:cpu_fa_tile
SongXiaoXi
SongXiaoXi ggml-cpu: enable tiled flash attention for non-vector-multiple head d…
4d5e0eac
SongXiaoXi SongXiaoXi requested a review from ggerganov ggerganov 2 days ago
github-actions github-actions added testing
github-actions github-actions added ggml
ggml-gh-bot
am17an
am17an commented on 2026-09-25
SongXiaoXi add AVX2 support for masked loading and storing in simd_gemm_ukernel_…
9257ffe5
am17an
ggml-gh-bot
SongXiaoXi
ggml-gh-bot

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone