[CUDA] Add DynamicSparseAttention (#32525)
### Description
Adds `com.microsoft.DynamicSparseAttention` v1 as a model-neutral CUDA
executor for externally selected attention indices.
- **Execution**
- `selected_only` over the contiguous main KV cache
- `local_plus_selected` over local main KV, auxiliary KV, and an
optional sink
- One numerically stable FP32 softmax across all participating entries
- **Attention preparation**
- Separate or packed QKV
- GQA/MQA head mapping
- QK RMSNorm and partial RoPE
- **Caching**
- Contiguous main KV append
- Past/present buffer aliasing
- Read-only auxiliary KV with optional shared K/V storage
- **Validation**
- Variable selected counts and `-1` padding
- Duplicate, bounds, and causal-index checks
- Mode/source and tensor-layout constraints
- **Integration**
- Contrib schema and shape inference
- CUDA registration and symbolic shape inference
- Operator documentation and deterministic CUDA coverage
```text
Qwen4-Exp: selected_only + main
DeepSeek V4: local_plus_selected + auxiliary
```
### Motivation and Context
Qwen4-Exp QSA and DeepSeek V4 CSA require a common attention executor
that consumes selector-produced indices without owning model-specific
scoring, TopK selection, or compression policy. This separates selection
policy from contiguous-cache attention while preserving DeepSeek’s
required joint local/selected/sink normalization.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: kunal-vaishnavi <115581922+kunal-vaishnavi@users.noreply.github.com>