Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
vllm-project/vllm
Pull Requests
Commits
Open
Closed
[CI][XPU] Skip CUDA-IPC weight sync metrics test on Intel
intel-gpu
ci/build
nvidia
#59556 opened 2026-10-01 06:42 by
zhenwei-intel
[Bugfix][Frontend] Check reused prompt token ids against the vocab before streaming
bug
ready
#59555 opened 2026-10-01 06:24 by
shijie-lyu
[Spec Decode][Watermarking] Support DFlash with dual_key_gumbel
documentation
speculative-decoding
needs-rebase
mrv2
dflash
#59554 opened 2026-10-01 06:13 by
shernshiou
[ROCm][Perf] Fuse the MXFP4 activation quant into the GDN gated RMSNorm
rocm
torch.compile
quantization
#59553 opened 2026-10-01 06:13 by
mjkvaak-amd
k3 chunk kda aiter enablment
kimi
k3
#59552 opened 2026-10-01 06:12 by
omuhamma
[ROCm][Bugfix] Fix ROCM_ATTN sliding-window boundary
bug
rocm
#59550 opened 2026-10-01 05:42 by
tangzzycc
[Rust Frontend] Support logprob_token_ids on OpenAI endpoints
rust
#59549 opened 2026-10-01 05:14 by
zupengwang
[KV Offload] Give each rank its own CPU region outside the replicated layout
#59547 opened 2026-10-01 04:35 by
Etelis
[KV Offload] Split the CPU offload region across shm files
#59546 opened 2026-10-01 04:35 by
Etelis
[KV Offload] Split replicated-layout pre-faulting across workers
#59545 opened 2026-10-01 04:35 by
Etelis
[KV Offload] Pin only each worker's own offload slots
#59544 opened 2026-10-01 04:35 by
Etelis
[Core] Opt-in per-request return of the last hidden states in generation (MRV2)
documentation
mrv2
scheduler
#59543 opened 2026-10-01 04:29 by
aminry
[Bugfix][Spec Decode] Keep heterogeneous-vocab draft models on Model Runner V1
bug
#59541 opened 2026-10-01 04:27 by
nishantrevur
[Core] Take the top-k/top-p Triton scratch from WorkspaceManager
mrv2
#59540 opened 2026-10-01 04:20 by
aoshen02
[Core] Reuse one CUDA graph capture stream per device
nvidia
#59537 opened 2026-10-01 04:15 by
aoshen02
[Bugfix][MRV2] Keep GDN prefill checkpoint metadata local to each cache group
bug
ready
mrv2
verified
#59536 opened 2026-10-01 04:14 by
ai-jz
[Core] Move FlashInfer MLA decode workspaces into WorkspaceManager
deepseek
nvidia
DSv4
DSv4.1
#59535 opened 2026-10-01 04:09 by
aoshen02
[Perf][Qwen4Exp] Merge QSA QKVG and indexer QK projections
qwen
#59533 opened 2026-10-01 03:19 by
gau-nernst
[Perf][DSv4.1] Run decoder replay layers in CUDA graphs and trim in PIECEWISE graphs
ready
torch.compile
deepseek
nvidia
mrv2
DSv4.1
#59532 opened 2026-10-01 03:09 by
Juntian777
[ROCm] Add gfx1151 W8A8 wvSplitK blockscale skinny path
performance
rocm
ci/build
deepseek
DSv4
#59531 opened 2026-10-01 03:02 by
keneoneth
[Bugfix][Bench] Clean up synthetic video files and writers
bug
performance
#59529 opened 2026-10-01 02:18 by
Prudhvivuda
[Bugfix] Keep the GLM-5.3 kpool tail and Qwen4 QSA ring out of the null block
bug
needs-rebase
qwen
glm
#59528 opened 2026-10-01 01:55 by
ivanium
[Core] Release the checkpoint page cache after weight loading
#59527 opened 2026-10-01 01:42 by
gracehonv
[CI] Run pip-compile pre-commit hooks only when requirements change
ci/build
#59526 opened 2026-10-01 01:20 by
tahsintunan
[CI] Use vllm_runner in fusions_e2e conftest for reliable GPU cleanup
ready
torch.compile
#59525 opened 2026-10-01 01:06 by
divakar-amd
[Bugfix][NIXL] Roll back incomplete remote-agent handshakes
bug
kv-connector
#59524 opened 2026-10-01 01:00 by
xijiaat
[ROCm] Enable the cuMem CUDA-graph pool offload on ROCm
rocm
torch.compile
nvidia
mrv2
#59523 opened 2026-10-01 00:48 by
aoshen02
[Feature][KV Connector] Add selective Mooncake recurrent checkpoint leases
documentation
performance
kv-connector
mooncake
#59522 opened 2026-10-01 00:47 by
0z5a
[CI][Test] Fix flaky weight sync metrics multi-API-server test
needs-rebase
#59521 opened 2026-10-01 00:23 by
aarushjain29
[ROCm][CI] Run distributed vision tests in the test_areas Distributed…
rocm
ci/build
#59519 opened 2026-10-01 00:10 by
aarushjain29
Newer
Older