fix latent packing order and validation generation in Ideogram4 LoRA trainer
- patchify_latents packed the latent channels as (ae, p_h, p_w) but the model's
packed layout (defined by Ideogram4Pipeline._decode) is (p_h, p_w, ae); every
training step fed the frozen base channel-permuted latents (with bn stats
applied to the wrong channels), so trained LoRAs corrupted generations at
inference while the training loss looked healthy.
- log_validation: drop torch.autocast entirely (corrupts Ideogram4 outputs:
fp16 -> NaN, bf16 -> gray), cast any fp32 params of the live transformer
(fp32 peft adapters on the quantized base + biases flipped by bitsandbytes'
in-forward bias.data cast) to the inference dtype for generation and restore
trainable params after, and strip accelerate's mixed-precision forward
wrapper (which re-introduces autocast) for the duration of validation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>