transformers
9921402b - [`distributed`] Default MoE `ep_plans` to token dispatch (#49160)

Commit
3 days ago
[`distributed`] Default MoE `ep_plans` to token dispatch (#49160) * Add dense and expert device mesh views with a MeshManager `initialize_distributed_mesh` now builds two named views of the same ranks: `(pp, fsdp, tp)` for dense layers and `(pp, efsdp, ep)` for experts, both keeping size-one axes so callers select dimensions by name. `MeshManager` routes `ep`/`efsdp` lookups to the expert view and everything else to the dense view. `DistributedConfig` gains `ep_size` (defaults to `tp_size` when `enable_expert_parallel=True`) and `efsdp_size`, with size validation. Model execution is unchanged: expert sharding and FSDP still use the `tp` and `fsdp` axes, and loading rejects `ep_size != tp_size` until the all-to-all dispatcher lands. * cleaning * clean * update * clean doc * Use ep_size in the Kimi K2.5 expert parallel example * Enable expert parallelism in the Mega MoE example * warn once * Add dense and expert device mesh views with a MeshManager `initialize_distributed_mesh` now builds two named views of the same ranks: `(pp, fsdp, tp)` for dense layers and `(pp, efsdp, ep)` for experts, both keeping size-one axes so callers select dimensions by name. `MeshManager` routes `ep`/`efsdp` lookups to the expert view and everything else to the dense view. `DistributedConfig` gains `ep_size` (defaults to `tp_size` when `enable_expert_parallel=True`) and `efsdp_size`, with size validation. Model execution is unchanged: expert sharding and FSDP still use the `tp` and `fsdp` axes, and loading rejects `ep_size != tp_size` until the all-to-all dispatcher lands. * cleaning * update * clean doc * Update src/transformers/distributed/mixin.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * style: ruff import block formatting in distributed config * add more comments * Decouple tp_plan and ep_plan for expert parallelism. Keep TP and EP plans separate on the model, resolve overrides via resolve_parallel_plans, and apply both through tensor parallel sharding on the tp mesh when ep_size matches tp_size. * Match expert paths with a regex and keep only expert rules under token dispatch Replace the fnmatch check in resolve_parallel_plans with a plan-pattern regex so plan keys are matched literally except for `*`. When the EP plan uses `ep_dispatch_experts`, drop the router masking rules from the EP plan and only take the expert modules and their parameters out of the TP plan. * revert change on warning * Apply suggestion from @ArthurZucker Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * cleaning * Fix tied embeddings for models with an EP-only base plan * [distributed] Add expert-parallel token dispatch, default for Qwen3 MoE * clean * fix * make check repo * Drop the load-time ep_size == tp_size guard that rejects token dispatch The check ran before the expert plan was resolved, so it rejected every token-dispatch layout (ep_size != tp_size). DistributedConfig._validate_resolved_ep_plan already enforces ep_size == tp_size for the masked plan once the plan is known. * remove all to all warning by using functional version * use ep plan instead of group_gemm is_expert * use dtensor instead of manuall all_reduce * better comment * [`distributed`]: Fix DeepSeek-V4 tied embeddings with combined TP and EP (#48824) Fix tied embeddings for models with an EP-only base plan * linting * Add dense and expert device mesh views with a MeshManager `initialize_distributed_mesh` now builds two named views of the same ranks: `(pp, fsdp, tp)` for dense layers and `(pp, efsdp, ep)` for experts, both keeping size-one axes so callers select dimensions by name. `MeshManager` routes `ep`/`efsdp` lookups to the expert view and everything else to the dense view. `DistributedConfig` gains `ep_size` (defaults to `tp_size` when `enable_expert_parallel=True`) and `efsdp_size`, with size validation. Model execution is unchanged: expert sharding and FSDP still use the `tp` and `fsdp` axes, and loading rejects `ep_size != tp_size` until the all-to-all dispatcher lands. * cleaning * clean * update * clean doc * Use ep_size in the Kimi K2.5 expert parallel example * Enable expert parallelism in the Mega MoE example * warn once * Add dense and expert device mesh views with a MeshManager `initialize_distributed_mesh` now builds two named views of the same ranks: `(pp, fsdp, tp)` for dense layers and `(pp, efsdp, ep)` for experts, both keeping size-one axes so callers select dimensions by name. `MeshManager` routes `ep`/`efsdp` lookups to the expert view and everything else to the dense view. `DistributedConfig` gains `ep_size` (defaults to `tp_size` when `enable_expert_parallel=True`) and `efsdp_size`, with size validation. Model execution is unchanged: expert sharding and FSDP still use the `tp` and `fsdp` axes, and loading rejects `ep_size != tp_size` until the all-to-all dispatcher lands. * cleaning * update * clean doc * Update src/transformers/distributed/mixin.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * style: ruff import block formatting in distributed config * add more comments * Decouple tp_plan and ep_plan for expert parallelism. Keep TP and EP plans separate on the model, resolve overrides via resolve_parallel_plans, and apply both through tensor parallel sharding on the tp mesh when ep_size matches tp_size. * Match expert paths with a regex and keep only expert rules under token dispatch Replace the fnmatch check in resolve_parallel_plans with a plan-pattern regex so plan keys are matched literally except for `*`. When the EP plan uses `ep_dispatch_experts`, drop the router masking rules from the EP plan and only take the expert modules and their parameters out of the TP plan. * revert change on warning * Apply suggestion from @ArthurZucker Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * cleaning * Fix tied embeddings for models with an EP-only base plan * [distributed] Add expert-parallel token dispatch, default for Qwen3 MoE * clean * fix * make check repo * Drop the load-time ep_size == tp_size guard that rejects token dispatch The check ran before the expert plan was resolved, so it rejected every token-dispatch layout (ep_size != tp_size). DistributedConfig._validate_resolved_ep_plan already enforces ep_size == tp_size for the masked plan once the plan is known. * remove all to all warning by using functional version * use ep plan instead of group_gemm is_expert * use dtensor instead of manuall all_reduce * better comment * linting * Train models sharded at load time with the Trainer, token dispatch included A model wrapped by its own FSDP2 (`DistributedConfig` with `fsdp_size > 1` or expert-parallel token dispatch) already owns placement and gradient reduction. Prepare it with `accelerator.prepare_model(..., evaluation_mode=True)` so Accelerate applies autocast and compilation without wrapping the DTensor parameters in DDP, which it rejects, or sharding them again. This is what token dispatch with `tp_size=1` needed: there is no `ParallelismConfig` to tell Accelerate about the model's TP, so it fell into the DDP branch. Loss scaling and token counting derive the replication factor from `get_tp_size()`, which now also falls back to Accelerate's parallelism config for tensor parallelism configured outside `DistributedConfig`. `tests/trainer/distributed/test_trainer_distributed_expert_parallel.py` trains a tiny Qwen3 MoE for four steps under masking on a `(fsdp, tp)` mesh, token dispatch with a batch per rank, and token dispatch with TP pairs, each with the same global batch as a single-process run, and compares the logged losses and gradient norms and the saved weights. * Use ep-dispatch version of expert_parallelism.md * Sync replicated trainable parameters when training a model sharded at load time A PEFT adapter attached after fully_shard is a plain tensor next to DTensor base weights. FSDP2 only reduces what it sharded and the DDP wrap is skipped on this path, so each rank trained its own adapter. Broadcast these parameters from rank 0 and average their gradient at the end of each accumulation window. * Fix utf-8 encoding in expert parallel trainer tests * [distributed] Default MoE expert-parallel plans to token dispatch Switch `base_model_ep_plan` from router masking with all-reduce (`ep_router` + `moe_tp_experts`) to all-to-all token dispatch (`ep_dispatch_experts`) for 41 MoE models, as done for Qwen3 MoE. Dispatch finds each expert's owner from the global expert ids, so the router rule is dropped. Models kept on masking: - Llama 4: no experts rule, non-standard experts module. - gemma4 (and diffusion_gemma), granitemoe_swa, hy_v4, inkling, mimo_v2_flash, youtu, zaya: their EP mixin tests already fail with masking because TP-sharded parameters used outside a TP style (attention sinks, norm weights, per-layer embeddings) mix Tensor and DTensor once TP composes with FSDP2, which dispatch always applies. Switch them once that composition is fixed. * remove * fix ruff * use attribute instead * Fix EP plans for youtu (dense, drop inherited plan) and zaya (token dispatch) * fix * revert * revert * typo * clearer * clean * cleaning * cleaning * dont torch cat for empty tokens * check code quality * Update src/transformers/distributed/utils.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update src/transformers/distributed/configuration_utils.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * linting * fix comment --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>
Author
Parents
Loading