llama.cpp
4b48a53b - server : optimize get_token_probabilities (#24796)

Commit
30 days ago
server : optimize get_token_probabilities (#24796) Use std::partial_sort to order only the requested top-n tokens instead of the full vocabulary logprobs sort: vocab=128000 n_top=0 iters=100 full sort: 8555.6 us/op partial sort: 704.3 us/op Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Author
Parents
Loading