transformers
89c57aff - Masked EP as its own parallel style, resolved from `ep_dispatch_experts` when `ep_size == tp_size`

Commit
2 days ago
Masked EP as its own parallel style, resolved from `ep_dispatch_experts` when `ep_size == tp_size` `EpMaskedExpertsParallel` extends `MoeExpertsParallel`, which already sums the partial outputs, the replicated inputs' gradients and handles local parameters; it adds only the routing: this rank's expert ids made local, the other ranks' routes moved past its experts with weight 0. Plan resolution swaps `ep_dispatch_experts` for it when the EP group shares one batch, and the mixin applies it like the router-masked plan. `EpDispatchExpertsParallel` is back to main's, refusing only experts that route themselves. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Parents
Loading