Add Granite-swa and Granitemoe-swa model support (#47179)
* Add GraniteSWA model: Granite with Sliding Window Attention + learnable attention sinks
- New model type 'granite_swa' with per-layer sliding window attention
- LSE-based sink mechanism: sink_scale = sigmoid(lse - sinks) applied post-attention
- Supports MuP parameters (embedding_multiplier, residual_multiplier, logits_scaling, attention_multiplier)
- Per-layer attention pattern: configurable full_attention vs sliding_attention layers
- Includes standalone modeling, configuration, modular source, and auto registration
* Add FA3 attention dispatch with native LSE for sink scaling
- Uses flash_attn_interface._flash_attn_forward for FA3 path (returns LSE natively)
- Falls back to eager attention with manual LSE computation
- Adds _supports_flash_attn and _compatible_flash_implementations flags
- Verified: both eager and FA3 produce correct output
* Enable DynamicSlidingWindowCache for memory-efficient KV cache
- SWA layers store only last (sliding_window - 1) tokens in KV cache
- Full attention layers keep complete KV history
- ~98% KV memory savings on SWA layers for long sequences
- Verified: identical results with and without sliding window cache
* Finalize GraniteSWA: modular impl, FA4 support, formatting, tests, docs
Make modular_granite_swa.py the single source of truth and regenerate the
modeling/config:
- Define GraniteSWAConfig in the modular file (inheriting GraniteConfig) so the converter no longer wipes the config, and rope_parameters standardizes correctly.
- Apply the learnable attention sink through the standard `s_aux` attention dispatch (as in GPT-OSS) instead of the bespoke `_lse`/`_fa3_lse` path. Eager keeps the original sigmoid(logsumexp - sink) math; FlashAttention-3 and 4 now work via the shared interface (`s_aux`/`learnable_sink`).
- Disable SDPA and FlexAttention (sink not expressible there); supported backends are eager, flash_attention_3, flash_attention_4.
- Add tests/models/granite_swa (CausalLMModelTest framework) and the model doc + _toctree entry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Add GraniteMoeSWA (MoeShared+SWA)
Add GraniteMoeSWA, sublassing GraniteMoeShared, with SWA sliding window + sink added.
- MoeSWAConfig inherits MoeSharedConfig, including default no shared experts (`shared_intermediate_size=0`).
- Attention uses SWAAttention and thus supports Flash3+4 and eager. No SDPA or Flex.
- MoeSWAModel handles per-layer causal masks during forward. All other layers inherit unchanged.
- Add tests/models/granitemoe_swa (CausalLMModelTest framework), model doc and _toctree entry, and register model in auto mappings.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Update model refs, add per-layer rope/nope support a la Llama4
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Small formatting/date fixes
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Test comments
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* automodel import style fix
* Regenerate modeling_ files post-resync to main
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Fix upload date in model cards
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Additional TP / historic test exclusions post-sync to main
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Docs cleanup, one final round on submit date
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Fix grante_swa init and imports
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Addressing review comments:
-Fix header years
-TP and integration tests
-Mark flex support
-Rope/Nope based on position embed passing
-Top level model defs via modular subclassing
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Replace binary rope/nope support with dynamic rope freq support
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Reorder router outputs to support native EP (and tests)
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Clean flex flags, flash correction tests, rope constructors
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Correct Flash2 test skip check
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Clean docstrings/comments, fix flash2 causal lm test
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Modelcard dates
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Fix rope typing and slow test shape checking
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Update integration test expectations by device
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* Regenerate modeling files to match inherited modular upstream changes
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
* adjust to ci
* fixup repo
* try
* kernels sync
---------
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
Co-authored-by: Bharat-Runwal <bharatrunwal@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: vasqu <antonprogamer@gmail.com>
Co-authored-by: Anton Vlasjuk <73884904+vasqu@users.noreply.github.com>