vllm
perf: Avoid copying inputs_embeds tensors to GPU unless prompt_embeds is enabled
#25739
Merged

perf: Avoid copying inputs_embeds tensors to GPU unless prompt_embeds is enabled #25739

qthequartermasterman
qthequartermasterman perf: Avoid copying inputs_embeds tensors to GPU unless prompt_embeds…
b3e5dc63
qthequartermasterman qthequartermasterman requested a review from WoosukKwon WoosukKwon 310 days ago
qthequartermasterman qthequartermasterman requested a review from robertgshaw2-redhat robertgshaw2-redhat 310 days ago
qthequartermasterman qthequartermasterman requested a review from njhill njhill 310 days ago
qthequartermasterman qthequartermasterman requested a review from ywang96 ywang96 310 days ago
qthequartermasterman qthequartermasterman requested a review from comaniac comaniac 310 days ago
qthequartermasterman qthequartermasterman requested a review from alexm-redhat alexm-redhat 310 days ago
mergify mergify added v1
gemini-code-assist
gemini-code-assist commented on 2025-09-26
DarkLight1337
qthequartermasterman fix: do not copy token ids from input batch to CPU tensor unless prom…
af40fa66
DarkLight1337
DarkLight1337 approved these changes on 2025-09-26
DarkLight1337 DarkLight1337 enabled auto-merge (squash) 310 days ago
github-actions github-actions added ready
vllm-bot vllm-bot merged d48f4d6d into main 310 days ago

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone