llama.cpp
3932b244
- WIP OCP FP8 E4M3 support
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
5 days ago
WIP OCP FP8 E4M3 support
References
#27512 - Quant: OCP FP8 E4M3 support
#28392 - Add FP8 KV-cache support
#28413 - CUDA: Add OCP FP8 support
#28898 - ggml: support quant scales for fp8 and nvfp4
#29791 - vulkan: add fp8 and scaled matmul support
Author
ORippler
Committer
ORippler
Parents
dcd387a4
Loading