Megatron checkpointing (#6293)
* Add bart fairseq run script
* Add frontend change to enable megatron
* Initial changes for checkpointing
* Megatron optim state loading, checkpoint aggregation, frontend distributed tests for H, D+H
* Add load_checkpoint changes
* Fix CI
* Cleanup
* Fix CI
* review comments
* review comments
* review comments: