onnxruntime
Optimize CUDA DynamicSparseAttention for large-context inference
#32671
Open

Optimize CUDA DynamicSparseAttention for large-context inference #32671

copilot-swe-agent
copilot-swe-agent copilot-swe-agent assigned copilot-swe-agent copilot-swe-agent 2 days ago
copilot-swe-agent copilot-swe-agent assigned kunal-vaishnavi kunal-vaishnavi 2 days ago
azure-pipelines
kunal-vaishnavi kunal-vaishnavi changed the base branch from main to copilot/copilotadd-cuda-dynamicsparseattention 2 days ago
Copilot Parallelize CUDA DynamicSparseAttention
cc469b31
Copilot Harden fused attention synchronization
f7e31683
Copilot Scale DynamicSparseAttention to large contexts
9db25a2e
Copilot Reduce large-context attention workspace
f4fde1bc
Copilot Bound sparse attention split workspace
0baf5060
Copilot Bound DynamicSparseAttention split work
e9c1dd69
kunal-vaishnavi kunal-vaishnavi force pushed from f9bc5f91 to e9c1dd69 2 days ago
xadupre xadupre requested a review from copilot-pull-request-reviewer copilot-pull-request-reviewer 2 days ago
copilot-pull-request-reviewer
copilot-pull-request-reviewer commented on 2026-09-17
Copilot Make DSA device validation opt-in
a755df8a
Copilot Test default DSA validation mode
2778f193
copilot-swe-agent copilot-swe-agent requested a review from kunal-vaishnavi kunal-vaishnavi 2 days ago

Login to write a write a comment.

Login via GitHub

Labels
Milestone