transformers
69a7fb1a - Add offload to gradient checkpointing (#48444)

Commit
27 days ago
Add offload to gradient checkpointing (#48444) * Add offload to gradient checkpointing Gradient checkpointing keeps one activation per checkpointed layer on the device, which is `layers x sequence x hidden` bytes and dominates at long sequence lengths. `offload` holds those in pinned host memory instead, through torch's `save_on_cpu`, so it follows whichever checkpointing path the model already uses. * Follow the repo conventions for the new option Document `offload` in `TrainingArguments.gradient_checkpointing_kwargs`, next to `every_n_layers`, and use the `backend_*` test helpers instead of `torch.accelerator` directly.
Author
Parents
Loading