llama.cpp
d34633d8 - clip : support more quantization types (#4846)

Commit

2 years ago

clip : support more quantization types (#4846) Uses ggml functions instead of hardcoded names and adds support to quantize into the modern Q-K variants. This is just the bare minimum to get k-types working - a more refined choice of types would be needed to get best quality on low quantizations. I ran a few tests, it doesn't break anything I could notice and a Q6_K ViT works almost as well as Q8_0 but 3 times the inference speed.

References

#4846 - clip.cpp quantization update

Author

cmp-nct

Parents

4f56458d

llama.cpp d34633d8 - clip : support more quantization types (#4846)

llama.cpp
d34633d8 - clip : support more quantization types (#4846)