onnxruntime
03ef37f5 - [CUDA] Add DynamicSparseAttention (#32525)

Commit
9 days ago
[CUDA] Add DynamicSparseAttention (#32525) ### Description Adds `com.microsoft.DynamicSparseAttention` v1 as a model-neutral CUDA executor for externally selected attention indices. - **Execution** - `selected_only` over the contiguous main KV cache - `local_plus_selected` over local main KV, auxiliary KV, and an optional sink - One numerically stable FP32 softmax across all participating entries - **Attention preparation** - Separate or packed QKV - GQA/MQA head mapping - QK RMSNorm and partial RoPE - **Caching** - Contiguous main KV append - Past/present buffer aliasing - Read-only auxiliary KV with optional shared K/V storage - **Validation** - Variable selected counts and `-1` padding - Duplicate, bounds, and causal-index checks - Mode/source and tensor-layout constraints - **Integration** - Contrib schema and shape inference - CUDA registration and symbolic shape inference - Operator documentation and deterministic CUDA coverage ```text Qwen4-Exp: selected_only + main DeepSeek V4: local_plus_selected + auxiliary ``` ### Motivation and Context Qwen4-Exp QSA and DeepSeek V4 CSA require a common attention executor that consumes selector-produced indices without owning model-specific scoring, TopK selection, or compression policy. This separates selection policy from contiguous-cache attention while preserving DeepSeek’s required joint local/selected/sink normalization. --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: kunal-vaishnavi <115581922+kunal-vaishnavi@users.noreply.github.com>
Author
Parents
Loading