llama.cpp
CUDA/HIP: Flash Attention tuning (gfx1201)
#28102
Merged

CUDA/HIP: Flash Attention tuning (gfx1201) #28102

pwilkin
pwilkin pwilkin requested a review from ggerganov ggerganov 18 days ago
pwilkin pwilkin requested a review 18 days ago
github-actions github-actions added testing
github-actions github-actions added ggml
github-actions github-actions added CUDA
Geramy
IMbackK IMbackK assigned IMbackK IMbackK 18 days ago
pwilkin
pwilkin pwilkin force pushed to d68f8769 17 days ago
IMbackK
IMbackK approved these changes on 2026-09-03
sbstnh
pwilkin
JohannesGaessler
JohannesGaessler requested changes on 2026-09-05
JohannesGaessler JohannesGaessler assigned JohannesGaessler JohannesGaessler 13 days ago
pwilkin
pwilkin
JohannesGaessler
pwilkin
JohannesGaessler
IMbackK
IMbackK
pwilkin
IMbackK
JohannesGaessler
JohannesGaessler commented on 2026-09-06
pwilkin pwilkin force pushed from f6264d8d to a3b28f93 12 days ago
pwilkin
JohannesGaessler
JohannesGaessler commented on 2026-09-06
JohannesGaessler
JohannesGaessler
eugene-su
Shardss
JohannesGaessler
pwilkin
gaperton
JohannesGaessler
pwilkin HIP: enable mma FA for head size 256 on RDNA4, tune configs
47e51920
pwilkin HIP: prefer whole-tile FA grids over stream-k on AMD WMMA
d2bf8ad7
JohannesGaessler revise stream_k logic
cbefddc5
JohannesGaessler revise kernel selection logic
461f7c1e
JohannesGaessler JohannesGaessler force pushed from 1873e6e3 to 461f7c1e 9 days ago
JohannesGaessler
JohannesGaessler
JohannesGaessler approved these changes on 2026-09-09
pwilkin
IMbackK
IMbackK approved these changes on 2026-09-11
IMbackK IMbackK merged 16378d93 into master 7 days ago

Login to write a write a comment.

Login via GitHub

Labels
Milestone