llama.cpp
07a19e27 - CUDA: fix quantized KV cache + multiple sequences (#14822)

Commit
46 days ago
CUDA: fix quantized KV cache + multiple sequences (#14822) * CUDA: fix quantized KV cache + multiple sequences * Update ggml/src/ggml-cuda/fattn-common.cuh Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Committer
Parents
Loading