transformers
d24918c4 - Add Granite-swa and Granitemoe-swa model support (#47179)

Commit
71 days ago
Add Granite-swa and Granitemoe-swa model support (#47179) * Add GraniteSWA model: Granite with Sliding Window Attention + learnable attention sinks - New model type 'granite_swa' with per-layer sliding window attention - LSE-based sink mechanism: sink_scale = sigmoid(lse - sinks) applied post-attention - Supports MuP parameters (embedding_multiplier, residual_multiplier, logits_scaling, attention_multiplier) - Per-layer attention pattern: configurable full_attention vs sliding_attention layers - Includes standalone modeling, configuration, modular source, and auto registration * Add FA3 attention dispatch with native LSE for sink scaling - Uses flash_attn_interface._flash_attn_forward for FA3 path (returns LSE natively) - Falls back to eager attention with manual LSE computation - Adds _supports_flash_attn and _compatible_flash_implementations flags - Verified: both eager and FA3 produce correct output * Enable DynamicSlidingWindowCache for memory-efficient KV cache - SWA layers store only last (sliding_window - 1) tokens in KV cache - Full attention layers keep complete KV history - ~98% KV memory savings on SWA layers for long sequences - Verified: identical results with and without sliding window cache * Finalize GraniteSWA: modular impl, FA4 support, formatting, tests, docs Make modular_granite_swa.py the single source of truth and regenerate the modeling/config: - Define GraniteSWAConfig in the modular file (inheriting GraniteConfig) so the converter no longer wipes the config, and rope_parameters standardizes correctly. - Apply the learnable attention sink through the standard `s_aux` attention dispatch (as in GPT-OSS) instead of the bespoke `_lse`/`_fa3_lse` path. Eager keeps the original sigmoid(logsumexp - sink) math; FlashAttention-3 and 4 now work via the shared interface (`s_aux`/`learnable_sink`). - Disable SDPA and FlexAttention (sink not expressible there); supported backends are eager, flash_attention_3, flash_attention_4. - Add tests/models/granite_swa (CausalLMModelTest framework) and the model doc + _toctree entry. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Add GraniteMoeSWA (MoeShared+SWA) Add GraniteMoeSWA, sublassing GraniteMoeShared, with SWA sliding window + sink added. - MoeSWAConfig inherits MoeSharedConfig, including default no shared experts (`shared_intermediate_size=0`). - Attention uses SWAAttention and thus supports Flash3+4 and eager. No SDPA or Flex. - MoeSWAModel handles per-layer causal masks during forward. All other layers inherit unchanged. - Add tests/models/granitemoe_swa (CausalLMModelTest framework), model doc and _toctree entry, and register model in auto mappings. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Update model refs, add per-layer rope/nope support a la Llama4 Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Small formatting/date fixes Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Test comments Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * automodel import style fix * Regenerate modeling_ files post-resync to main Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Fix upload date in model cards Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Additional TP / historic test exclusions post-sync to main Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Docs cleanup, one final round on submit date Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Fix grante_swa init and imports Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Addressing review comments: -Fix header years -TP and integration tests -Mark flex support -Rope/Nope based on position embed passing -Top level model defs via modular subclassing Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Replace binary rope/nope support with dynamic rope freq support Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Reorder router outputs to support native EP (and tests) Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Clean flex flags, flash correction tests, rope constructors Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Correct Flash2 test skip check Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Clean docstrings/comments, fix flash2 causal lm test Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Modelcard dates Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Fix rope typing and slow test shape checking Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Update integration test expectations by device Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * Regenerate modeling files to match inherited modular upstream changes Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> * adjust to ci * fixup repo * try * kernels sync --------- Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> Co-authored-by: Bharat-Runwal <bharatrunwal@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: vasqu <antonprogamer@gmail.com> Co-authored-by: Anton Vlasjuk <73884904+vasqu@users.noreply.github.com>
Author
Parents
Loading