DeepSpeed
9311fd55 - Route isend/irecv to nonblocking backend methods and stage them as async on MPS (#8303)

Commit
1 day ago
Route isend/irecv to nonblocking backend methods and stage them as async on MPS (#8303) ## Summary Resolves @delock's review note on #8293 (https://github.com/deepspeedai/DeepSpeed/pull/8293#discussion_r3838066086): `irecv` is asynchronous by contract and has no `async_op` parameter, so the MPS CPU-staging wrapper must handle it explicitly. Two fixes: 1. **`deepspeed/comm/comm.py`** — `isend`/`irecv` dispatched to the *blocking* `cdb.send`/`cdb.recv` (since the original comm backend, #1985). Callers got a blocking call and `recv`'s return value (the source rank `int`) instead of a waitable handle, so `dist.irecv(...).wait()` raised `AttributeError`. This affects every backend, not just MPS — e.g. the 1-bit comm helpers (`runtime/comm/{compressed,hccl,nccl}.py`) call `dist.isend/irecv(...).wait()`. They now route to `cdb.isend`/`cdb.irecv`. 2. **`deepspeed/comm/torch.py`** — with the routing fixed, the MPS staging wrapper's copy-back decision (keyed on an `async_op` argument) ran immediately for `irecv`, before the transfer completed. A new `always_async` flag on `stage_on_cpu` defers the copy-back to the handle's `wait()` for `isend`/`irecv`. `StagedWork.wait()` now also returns the underlying work's wait result. ### Verified (M5 Max, macOS 26.3, torch 2.13) - Real two-process gloo run with MPS tensors: on master, `dist.irecv` returns an `int` and `.wait()` crashes; with this PR it returns a handle and the buffer holds the correct payload after `wait()`. - `DS_ACCELERATOR=mps pytest unit/comm/test_dist.py`: 10 passed (multi-rank cases skip on 1 device). - ZeRO-2/3 smoke training unaffected. ### Test Adds `TestDistIsendIrecv` (world size 2) to the existing `tests/unit/comm/test_dist.py`: rank 0 `isend`s, rank 1 `irecv`s, both assert a waitable handle and verify the payload after `wait()`. Backend-agnostic, so it exercises the routing fix on CUDA/CPU CI as well. ### Relation to #8301 #8301 addresses the same note with a more extensive `StagedWork` (futures, result identity restoration, weakref buffer tracking). This PR makes the fix more concise and accurate: no current DeepSpeed users calls `Work.result()`/`get_future()` on staged P2P ops, and the staged CPU buffer for `isend` is kept alive by the deferred copy-back closure until `wait()`. Huge Credit to @FU-max-boop for the thorough analysis of the Work semantics and fixes. --------- Signed-off-by: PKUWZP <zhipeng.rainbowserie@gmail.com>
Author
Parents
Loading