Add offload to gradient checkpointing (#48444)
* Add offload to gradient checkpointing
Gradient checkpointing keeps one activation per checkpointed layer on the device, which is
`layers x sequence x hidden` bytes and dominates at long sequence lengths. `offload` holds those in
pinned host memory instead, through torch's `save_on_cpu`, so it follows whichever checkpointing path
the model already uses.
* Follow the repo conventions for the new option
Document `offload` in `TrainingArguments.gradient_checkpointing_kwargs`, next to `every_n_layers`,
and use the `backend_*` test helpers instead of `torch.accelerator` directly.