transformers
28cb4827 - Declare dispatch_experts_forward next to its only caller

Commit
30 days ago
Declare dispatch_experts_forward next to its only caller `dispatch_experts_forward` and the `_ScaleGrad` helper it uses move from `integrations/moe.py` to `distributed/tensor_parallel.py`, above `EpDispatchExpertsParallel`, so the local import inside `install_forward` goes away. Also say in the FSDP2 section that it describes expert parallelism without token dispatch, and what changes when dispatch is on, which read as a contradiction with the section above it.
Author
Parents
Loading