Fix silent gradient-bucket drop on CPU in ZeRO-1/2 non-contiguous reduction (#8382)
## Problem
While investigating multi-rank CPU CI (deepspeedai/DeepSpeed#8381),
`TestGradientAllreduceOp` failed for two variants (`sum` +
`prescale_gradients=True` + ZeRO-1/2 + `contiguous_gradients=False` +
`reduce_scatter=False`) with results off by a factor of world_size, and
`TestUnmanagedGradientAccumulationOffload` showed a numeric mismatch.
Separately, any test comparing tensors across ranks on CPU crashed with
`RuntimeError: Invalid device string` or a gloo allgather error.
## Root cause
`split_half_float_double()` built legacy type strings from the
accelerator device name (`torch.{device}.FloatTensor`). CPU tensors
report `torch.FloatTensor` with **no device prefix**, so no tensor ever
matched: the bucket list came back empty and **the gradient all-reduce
was silently skipped** — every rank updated on its local gradients.
- With `gradient_allreduce_op=mean` this stays invisible: identical
per-rank gradients make the unreduced local value equal the mean.
- With `sum` the result is wrong by a factor of world_size.
The same variants pass on GPU CI (all 18 `TestGradientAllreduceOp`
variants and 5 `TestGradientAllreduceOpTraining` variants are green on
`modal-torch-latest`), because `torch.cuda.FloatTensor` matches on CUDA
— the bug is CPU-specific, and CPU multi-rank runs were the first to
reach this code.
The harness crashes were two stacked bugs in `reduce_boolean_flags()`
(tests/unit/common.py): `current_device()` returns a bare rank index (a
`LOCAL_RANK` string on CPU) that torch cannot parse as a device, and
gloo rejects 0-dim inputs to `all_gather_into_tensor`.
## Fix
- Group gradient buckets by `t.dtype` instead of legacy type strings,
making bucketing device-independent (behavior-equivalent on CUDA, where
the legacy strings matched).
- Use `current_device_name()` (valid on every accelerator) and carry the
flag as a 1-element tensor.
## Verification (CPU/gloo, world sizes 2–3)
| Test | Before | After |
|---|---|---|
| `TestGradientAllreduceOp` (18 variants) | 2 failed | 18/18 pass |
| `TestGradientAllreduceOpTraining` (5) | crash in comparison helper |
pass (3 locally + muon skips fp16) |
| `TestUnmanagedGradientAccumulationOffload` + `InactiveParams` (6) |
crash + numeric mismatch | 6/6 pass |
| `TestZero3ParamPartitioningBase` / `TestZeroToFP32` (regression smoke)
| — | pass |
GPU behavior is unchanged by construction (same buckets, same order);
GPU CI will re-confirm.
## Follow-up
With the comparison helper fixed,
`test_unmanaged_varying_backward_count[3]` (ZeRO-3) can now reach its
numeric comparison and reveals a pre-existing unmanaged-vs-managed
mismatch on CPU that was previously masked by the harness crash. Out of
scope here.
---------
Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Signed-off-by: Ma, Guokai <guokai.ma@intel.com>