Trainer: broadcast the parameters FSDP2 does not shard before training
SFTTrainer attaches the PEFT adapter before Trainer.__init__ seeds the RNG, so each rank
initialises lora_A from its own random state. DDP hid this by broadcasting rank 0's
parameters at construction; the native-FSDP2 path skips DDP. Broadcast them here, next to
the gradient averaging, so the replicas start equal and stay equal.
(cherry picked from commit 3961edd9b1b2279117e3f0136ce69e66e75f4460)