llama.cpp
69e62fc7 - llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized (#25871)

Commit
7 days ago
llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized (#25871) * llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized * llama : enforce the same K and V cache types for MLA models --------- Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Author
Parents
Loading