vllm
[SpecDecode] Reduce TP communication for large-vocab draft models speculative decoding
#39419
Merged

Loading