DeepSpeed
20a60f22 - Fix silent gradient-bucket drop on CPU in ZeRO-1/2 non-contiguous reduction (#8382)

Commit
13 days ago
Fix silent gradient-bucket drop on CPU in ZeRO-1/2 non-contiguous reduction (#8382) ## Problem While investigating multi-rank CPU CI (deepspeedai/DeepSpeed#8381), `TestGradientAllreduceOp` failed for two variants (`sum` + `prescale_gradients=True` + ZeRO-1/2 + `contiguous_gradients=False` + `reduce_scatter=False`) with results off by a factor of world_size, and `TestUnmanagedGradientAccumulationOffload` showed a numeric mismatch. Separately, any test comparing tensors across ranks on CPU crashed with `RuntimeError: Invalid device string` or a gloo allgather error. ## Root cause `split_half_float_double()` built legacy type strings from the accelerator device name (`torch.{device}.FloatTensor`). CPU tensors report `torch.FloatTensor` with **no device prefix**, so no tensor ever matched: the bucket list came back empty and **the gradient all-reduce was silently skipped** — every rank updated on its local gradients. - With `gradient_allreduce_op=mean` this stays invisible: identical per-rank gradients make the unreduced local value equal the mean. - With `sum` the result is wrong by a factor of world_size. The same variants pass on GPU CI (all 18 `TestGradientAllreduceOp` variants and 5 `TestGradientAllreduceOpTraining` variants are green on `modal-torch-latest`), because `torch.cuda.FloatTensor` matches on CUDA — the bug is CPU-specific, and CPU multi-rank runs were the first to reach this code. The harness crashes were two stacked bugs in `reduce_boolean_flags()` (tests/unit/common.py): `current_device()` returns a bare rank index (a `LOCAL_RANK` string on CPU) that torch cannot parse as a device, and gloo rejects 0-dim inputs to `all_gather_into_tensor`. ## Fix - Group gradient buckets by `t.dtype` instead of legacy type strings, making bucketing device-independent (behavior-equivalent on CUDA, where the legacy strings matched). - Use `current_device_name()` (valid on every accelerator) and carry the flag as a 1-element tensor. ## Verification (CPU/gloo, world sizes 2–3) | Test | Before | After | |---|---|---| | `TestGradientAllreduceOp` (18 variants) | 2 failed | 18/18 pass | | `TestGradientAllreduceOpTraining` (5) | crash in comparison helper | pass (3 locally + muon skips fp16) | | `TestUnmanagedGradientAccumulationOffload` + `InactiveParams` (6) | crash + numeric mismatch | 6/6 pass | | `TestZero3ParamPartitioningBase` / `TestZeroToFP32` (regression smoke) | — | pass | GPU behavior is unchanged by construction (same buckets, same order); GPU CI will re-confirm. ## Follow-up With the comparison helper fixed, `test_unmanaged_varying_backward_count[3]` (ZeRO-3) can now reach its numeric comparison and reveals a pre-existing unmanaged-vs-managed mismatch on CPU that was previously masked by the harness crash. Out of scope here. --------- Signed-off-by: Guokai Ma <guokai.ma@intel.com> Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
Author
Parents
Loading