llama.cpp
74c13eea - metal : dequantize q4_0, q4_1, q5_0 and q5_1 KV to f16 before flash attention

Commit
22 hours ago
metal : dequantize q4_0, q4_1, q5_0 and q5_1 KV to f16 before flash attention The dequant pass now covers all quantized KV types supported by the Metal flash attention kernels. The dequant kernel, kargs, scratch allocation and dispatch are type-generic, so each type is one kernel instantiation plus one gate case. Assisted-by: pi:llama.cpp/Qwen3.8-27B
Author
Parents
Loading