transformers
9d0eb661 - Trainer: broadcast the parameters FSDP2 does not shard before training

Commit
9 days ago
Trainer: broadcast the parameters FSDP2 does not shard before training SFTTrainer attaches the PEFT adapter before Trainer.__init__ seeds the RNG, so each rank initialises lora_A from its own random state. DDP hid this by broadcasting rank 0's parameters at construction; the native-FSDP2 path skips DDP. Broadcast them here, next to the gradient averaging, so the replicas start equal and stay equal. (cherry picked from commit 3961edd9b1b2279117e3f0136ce69e66e75f4460)
Author
Committer
Parents
Loading