llama.cpp
1d7db017
- Materialize input_scales also for FP8 when converting
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
4 days ago
Materialize input_scales also for FP8 when converting
References
#27512 - Quant: OCP FP8 E4M3 support
#28392 - Add FP8 KV-cache support
#28413 - CUDA: Add OCP FP8 support
#28898 - ggml: support quant scales for fp8 and nvfp4
#29791 - vulkan: add fp8 and scaled matmul support
Author
ORippler
Committer
ORippler
Parents
eff0a3db
Loading