onnxruntime
2e7de545 - [CUDA] Fix SparseAttention Kernel (#20716)

Commit

1 year ago

[CUDA] Fix SparseAttention Kernel (#20716) ### Description Currently, there is one bool flag to indicate whether kernel is loaded. However, there are v1 and v2 kernels, so the flag will allow only one version of kernel loaded. We use v1 kernel for prompt and v2 kernel for token generation, and the flag will cause issue when we want both prompt and token generation. This bug is found in integration test. The unit test only test one kernel at a time so the issue was not found before. Another possible walkaround without this fix is to set an environment variable `ORT_DISABLE_SPARSE_ATTENTION_V1=1` ### Motivation and Context

References

#20716 - [CUDA] Fix SparseAttention Kernel

Author

tianleiwu

Parents

d7f7c3b3

onnxruntime 2e7de545 - [CUDA] Fix SparseAttention Kernel (#20716)

onnxruntime
2e7de545 - [CUDA] Fix SparseAttention Kernel (#20716)