transformers
35924ec3 - Support per-layer cache configuration (#48178)

Commit
3 days ago
Support per-layer cache configuration (#48178) * Support heterogeneous cache layouts * Add heterogeneous cache compatibility tests * revert gpt-oss mxfp4 changes * Preserve wrapped forward metadata in Accelerate hooks * Revert "Preserve wrapped forward metadata in Accelerate hooks" This reverts commit 7937a3b64a5f9568455899f1e356b9dafd1f566c. * Remove underscore from _get_layer_types_and_kwargs * Remove get_representative_kv_layer_idx and restore get_seq_length * remove related tests * fix * Clean up masking utils changes * Fix doc strings * Fix per-layer static cache initialization and empty cache cropping * Fix per-layer head shape inference for static caches * clean up tests * Remove unrelated changes based on the review --------- Co-authored-by: Elad Segal <13485709+eladsegal@users.noreply.github.com>
Author
Parents
Loading