DeepSpeed
aa49eaab - fix(rollout): reject shared prefill with continuous batching (#8583)

Commit
9 days ago
fix(rollout): reject shared prefill with continuous batching (#8583) ## Why `HybridEngineRollout.generate` selects the continuous path whenever `SamplingConfig.continuous_batch_size` is set, before any shared-prefill logic runs. With `HybridEngineRolloutConfig.use_shared_prefill=True`, that flag was therefore ignored with no error. This PR rejects the combination explicitly, matching the approach already taken for CUDA graph capture on the continuous path. Fixes #8458. ## Scope - `HybridEngineRollout._validate_continuous_inputs`: raise `ValueError` when `use_shared_prefill` is enabled. - `docs/code-docs/source/inference-engine.rst`: document that shared prefill and continuous batching cannot be combined. Out of scope: implementing shared-prefill semantics inside continuous batching. ## Tradeoffs Rejection over a compatible shared-prefill continuous mode. Continuous batching already requires `n_samples_per_prompt=1`, so shared prefill has no clear benefit there today. ## Blast Radius Only the experimental continuous path changes, and only when shared prefill is also requested. Default `generate()` and shared-prefill-only configs are unchanged. ## Verification Before the fix, `_validate_continuous_inputs` accepted `use_shared_prefill=True` with `continuous_batch_size=1` (silent ignore). After the fix it raises `ValueError: continuous batching does not support shared prompt prefill`. ``` TRANSFORMERS_NO_TF=1 PYTHONPATH=. python3 -m pytest tests/unit/runtime/rollout/test_hybrid_engine_rollout.py -q # 32 passed, 1 skipped ``` Signed-off-by: Chandan Kumar <cml.codes@gmail.com>
Author
Parents
Loading