[c10d] Make reduce_scatter as a custom op (#79683)
Summary:
This patch makes reduce_scatter as a custom op such that it's dispatcher
passable. It's one part of the effort to route comm ops to the dispatcher
such that tracing mechanisms that relies on the dispatcher can trace them,
e.g., LazyTensor and AOTAutograd.
Test Plan:
python test/distributed/test_c10d_nccl.py -k test_reduce_scatter_ops
Pull Request resolved: https://github.com/pytorch/pytorch/pull/79683
Approved by: https://github.com/mrshenli