onnxruntime
[CUDA] Add TensorRT fused attention fp16 v2 kernels
#12814
Merged

[CUDA] Add TensorRT fused attention fp16 v2 kernels #12814

tianleiwu merged 15 commits into main from tlwu/trt_fused_attention
tianleiwu
tianleiwu Add TensorRT fused attention fp16 kernels
5ed7df75
tianleiwu tianleiwu marked this pull request as draft 3 years ago
tianleiwu tianleiwu changed the title [WIP] Add TensorRT fused attention fp16 v2 kernels [CUDA] Add TensorRT fused attention fp16 v2 kernels 3 years ago
tianleiwu fix cpplint warning
e04a2c04
tianleiwu dump half tensor
a7f0302f
tianleiwu update test thresholds
7203908f
tianleiwu drop sm 72; fix lint warnings
55540eab
tianleiwu sm86=>sm80 fall back; remove sm_72
2439f680
tianleiwu fix TRT kernel input format; exclude sm75_s512_d64
e1abe53e
tianleiwu rewrite AddBiasTransposeTrt
19ce6bb5
tianleiwu remove seq 512 for sm75; and head_size 32 kernels
96f71476
tianleiwu Add env variable ORT_DISABLE_FUSED_ATTENTION
3e7f3f81
tianleiwu exclude files in hipify
0416aac6
tianleiwu fix cpplint; AttentionPastState_dynamic threshold
f327fec0
tianleiwu format
5aa9e2e2
tianleiwu fix --use_mask_index in benchmark
b1440e56
tianleiwu Merge branch 'main' into tlwu/trt_fused_attention
52ff22e1
tianleiwu tianleiwu marked this pull request as ready for review 3 years ago
tianleiwu tianleiwu requested a review from gh-yewang gh-yewang 3 years ago
tianleiwu tianleiwu requested a review from hariharans29 hariharans29 3 years ago
tianleiwu tianleiwu requested a review from yufenglee yufenglee 3 years ago
gh-yewang
gh-yewang approved these changes on 2022-09-13
tianleiwu tianleiwu merged 95c4fc68 into main 3 years ago
tianleiwu tianleiwu deleted the tlwu/trt_fused_attention branch 3 years ago
philschmid
hariharans29
hariharans29 commented on 2022-09-20

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone