EanWang211123marked this pull request as ready for review 157 days ago
EanWang211123
changed the title [SpecDecode] Reduce TP communication for large-vocab draft models in DFlash/PARD [SpecDecode] Reduce TP communication for large-vocab draft models in DFlash/PARD speculative decoding157 days ago
EanWang211123
changed the title [SpecDecode] Reduce TP communication for large-vocab draft models in DFlash/PARD speculative decoding [SpecDecode] Reduce TP communication for large-vocab draft models speculative decoding105 days ago
[update] remove draft_id_to_target_id init logits in llama_eagle3
Login to write a write a comment.
Login via GitHub