benchmark
cb0f1fb8 - Reproduce Triton, sdpa-flash, and sdpa-cudnn attention

Commit
2 years ago
Reproduce Triton, sdpa-flash, and sdpa-cudnn attention Summary: We are open-sourcing the `flash-attention` operator so that NVIDIA can reproduce its results. Add label to `cudnn` backend. Reviewed By: jianyuh Differential Revision: D58418594 fbshipit-source-id: f530793e63c3023c5fc57c4dfe06319e18c58b57
Author
Parents
Loading