vllm
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity
#38479
Merged

Loading