llama : support quantum K cache (#4312)

Commit

2 years ago

llama : support quantum K cache (#4312) * llama : support quantum K cache (wip) * metal : add F32 -> Q8_0 copy kernel * cuda : add F32 -> Q8_0 copy kernel ggml-ci * cuda : use mmv kernel for quantum cache ops * llama : pass KV cache type through API * llama : fix build ggml-ci * metal : add F32 -> Q4_0 copy kernel * metal : add F32 -> Q4_1 copy kernel * cuda : wip * cuda : add F32 -> Q4_0 and F32 -> Q4_1 copy kernels * llama-bench : support type_k/type_v * metal : use mm kernel only for quantum KV cache * cuda : add comment * llama : remove memory_f16 and kv_f16 flags --------- Co-authored-by: slaren <slarengh@gmail.com>

References

#4312 - llama : support quantum K cache

Author

ggerganov

Parents

66aaac98

llama.cpp 1a1a1c38 - llama : support quantum K cache (#4312)

llama.cpp
1a1a1c38 - llama : support quantum K cache (#4312)