onnxruntime
Support Smooth Softmax in GroupQueryAttention
#21867
Merged
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Overview
Commits
9
Changes
View On
GitHub
Support Smooth Softmax in GroupQueryAttention
#21867
tianleiwu
merged 9 commits into
main
from
smooth_softmax_gqa_flash_and_cpu
Add smooth softmax to GQA cpu
d3f9caba
update gqa cuda benchmark for smooth_softmax
5551c6ec
update gqa cpu test for smooth softmax
7085894a
test gqa cuda smooth_softmax
8678ede6
flash attn smooth softmax
198accfb
MLAS smooth softmax
517bb086
not test smooth softmax on efficient attention
1d87fe22
Merge branch 'main' into smooth_softmax_gqa_flash_and_cpu
ea9d8c28
update op doc
7a8e1b6b
tianleiwu
requested a review
2 years ago
tianleiwu
requested a review
from
gh-yewang
2 years ago
tianleiwu
requested a review
from
yufenglee
2 years ago
yufenglee
approved these changes on 2024-08-27
kunal-vaishnavi
approved these changes on 2024-08-27
tianleiwu
merged
6e575769
into main
2 years ago
tianleiwu
deleted the smooth_softmax_gqa_flash_and_cpu branch
2 years ago
Login to write a write a comment.
Login via GitHub
Reviewers
kunal-vaishnavi
yufenglee
gh-yewang
Assignees
No one assigned
Labels
None yet
Milestone
No milestone
Login to write a write a comment.
Login via GitHub