llama.cpp
b3daa077 - vulkan: sparse flash attention for quantized K/V (#29639)

Commit
2 days ago
vulkan: sparse flash attention for quantized K/V (#29639) * vulkan: sparse flash attention for quantized K/V Assisted-by: Claude * vulkan: single-scan sparse FA index compaction The compaction ran one workgroup per mask row and walked the row in BLOCK_SIZE chunks, with a workgroup scan per chunk. For decode that is one workgroup doing KV/1024 barrier-bound iterations, so at 128k cells it cost more than the sparse attention it feeds. Split the row into contiguous segments instead: one per subgroup with ballot counting over coalesced loads, or one per thread without subgroups. A single scan over the segment counts then gives each segment its output offset. The index list stays ascending.
Author
Parents
Loading