Add SDPA support for PatchTST model (#42465)
* Add SDPA and Flash Attention support for PatchTST model
- Add _supports_sdpa = True and _supports_flash_attn = True to PatchTSTPreTrainedModel
- The existing PatchTSTAttention class already uses ALL_ATTENTION_FUNCTIONS
to select the attention implementation based on config._attn_implementation
- Fix test_modeling_patchtst.py _prepare_for_class for dynamic batch sizes
* Guard PatchTST positional init under ZeRO-3
* Force SDPA in PatchTST regression integration test
* Use sdpa attn in PatchTST regression test
* fixups re tests
---------
Co-authored-by: Kashif Rasul <kashif.rasul@gmail.com>
Co-authored-by: vasqu <antonprogamer@gmail.com>
Co-authored-by: Anton Vlasjuk <73884904+vasqu@users.noreply.github.com>