Support per-layer cache configuration (#48178)
* Support heterogeneous cache layouts
* Add heterogeneous cache compatibility tests
* revert gpt-oss mxfp4 changes
* Preserve wrapped forward metadata in Accelerate hooks
* Revert "Preserve wrapped forward metadata in Accelerate hooks"
This reverts commit 7937a3b64a5f9568455899f1e356b9dafd1f566c.
* Remove underscore from _get_layer_types_and_kwargs
* Remove get_representative_kv_layer_idx and restore get_seq_length
* remove related tests
* fix
* Clean up masking utils changes
* Fix doc strings
* Fix per-layer static cache initialization and empty cache cropping
* Fix per-layer head shape inference for static caches
* clean up tests
* Remove unrelated changes based on the review
---------
Co-authored-by: Elad Segal <13485709+eladsegal@users.noreply.github.com>