Route isend/irecv to nonblocking backend methods and stage them as async on MPS (#8303)
## Summary
Resolves @delock's review note on #8293
(https://github.com/deepspeedai/DeepSpeed/pull/8293#discussion_r3838066086):
`irecv` is asynchronous by contract and has no `async_op` parameter, so
the MPS CPU-staging wrapper must handle it explicitly.
Two fixes:
1. **`deepspeed/comm/comm.py`** — `isend`/`irecv` dispatched to the
*blocking* `cdb.send`/`cdb.recv` (since the original comm backend,
#1985). Callers got a blocking call and `recv`'s return value (the
source rank `int`) instead of a waitable handle, so
`dist.irecv(...).wait()` raised `AttributeError`. This affects every
backend, not just MPS — e.g. the 1-bit comm helpers
(`runtime/comm/{compressed,hccl,nccl}.py`) call
`dist.isend/irecv(...).wait()`. They now route to
`cdb.isend`/`cdb.irecv`.
2. **`deepspeed/comm/torch.py`** — with the routing fixed, the MPS
staging wrapper's copy-back decision (keyed on an `async_op` argument)
ran immediately for `irecv`, before the transfer completed. A new
`always_async` flag on `stage_on_cpu` defers the copy-back to the
handle's `wait()` for `isend`/`irecv`. `StagedWork.wait()` now also
returns the underlying work's wait result.
### Verified (M5 Max, macOS 26.3, torch 2.13)
- Real two-process gloo run with MPS tensors: on master, `dist.irecv`
returns an `int` and `.wait()` crashes; with this PR it returns a handle
and the buffer holds the correct payload after `wait()`.
- `DS_ACCELERATOR=mps pytest unit/comm/test_dist.py`: 10 passed
(multi-rank cases skip on 1 device).
- ZeRO-2/3 smoke training unaffected.
### Test
Adds `TestDistIsendIrecv` (world size 2) to the existing
`tests/unit/comm/test_dist.py`: rank 0 `isend`s, rank 1 `irecv`s, both
assert a waitable handle and verify the payload after `wait()`.
Backend-agnostic, so it exercises the routing fix on CUDA/CPU CI as
well.
### Relation to #8301
#8301 addresses the same note with a more extensive `StagedWork`
(futures, result identity restoration, weakref buffer tracking). This PR
makes the fix more concise and accurate: no current DeepSpeed users
calls `Work.result()`/`get_future()` on staged P2P ops, and the staged
CPU buffer for `isend` is kept alive by the deferred copy-back closure
until `wait()`. Huge Credit to @FU-max-boop for the thorough analysis of
the Work semantics and fixes.
---------
Signed-off-by: PKUWZP <zhipeng.rainbowserie@gmail.com>