Reproduce Triton, sdpa-flash, and sdpa-cudnn attention
Summary:
We are open-sourcing the `flash-attention` operator so that NVIDIA can reproduce its results.
Add label to `cudnn` backend.
Reviewed By: jianyuh
Differential Revision: D58418594
fbshipit-source-id: f530793e63c3023c5fc57c4dfe06319e18c58b57