llama.cpp
8afc93cd - tests : add perf cases for sparse flash attention prefill

Commit
3 days ago
tests : add perf cases for sparse flash attention prefill Measure the sparse vec FA kernel across KV sizes, n_kv_max hints and batch sizes. Run with: ./build/bin/test-backend-ops -b MTL0 -o FLASH_ATTN_EXT -p "n_kv_max=[1-9]" perf Assisted-by: pi:llama.cpp/DeepSeek-v4-0731
Author
Committer
Parents
Loading