llama.cpp
cc231cb0 - dflash: pass missing NVFP4 scales to attention operations (#28000)

Commit
3 days ago
dflash: pass missing NVFP4 scales to attention operations (#28000) - DFlash2 NVFP4 draft models produced almost no accepted speculative tokens because the Q, K, V, and output projection scales were not passed to the corresponding graph operations.
Author
Parents
Loading