llama.cpp
16378d93 - CUDA/HIP: Flash Attention tuning (gfx1201) (#28102)

Commit
3 days ago
CUDA/HIP: Flash Attention tuning (gfx1201) (#28102) * HIP: enable mma FA for head size 256 on RDNA4, tune configs Assisted-by: Claude Assisted-by: Codex * HIP: prefer whole-tile FA grids over stream-k on AMD WMMA Assisted-by: Claude Assisted-by: Codex * revise stream_k logic * revise kernel selection logic --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Author
Parents
Loading