llama.cpp
cc231cb0
- dflash: pass missing NVFP4 scales to attention operations (#28000)
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
3 days ago
dflash: pass missing NVFP4 scales to attention operations (#28000) - DFlash2 NVFP4 draft models produced almost no accepted speculative tokens because the Q, K, V, and output projection scales were not passed to the corresponding graph operations.
References
#28000 - dflash: pass missing NVFP4 scales to attention operations
Author
JamePeng
Parents
bebc9350
Loading