llama.cpp
f896d2c3 - server: improve speed of speculative decoding (#17808)

Commit
2 days ago
server: improve speed of speculative decoding (#17808) * server: improve speed of speculative decoding * fix small draft case * add link to the PR * server : fix generation time measurement * server : fix draft acceptance logs (add SRV_CNT, SLT_CNT macros) * server : add comment * add PR to docs --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Author
Parents
Loading