[CUDA] Add TensorRT fused attention fp16 v2 kernels #12814
Add TensorRT fused attention fp16 kernels
5ed7df75
tianleiwu
marked this pull request as draft 3 years ago
tianleiwu
changed the title [WIP] Add TensorRT fused attention fp16 v2 kernels [CUDA] Add TensorRT fused attention fp16 v2 kernels 3 years ago
fix cpplint warning
e04a2c04
dump half tensor
a7f0302f
update test thresholds
7203908f
drop sm 72; fix lint warnings
55540eab
sm86=>sm80 fall back; remove sm_72
2439f680
fix TRT kernel input format; exclude sm75_s512_d64
e1abe53e
rewrite AddBiasTransposeTrt
19ce6bb5
remove seq 512 for sm75; and head_size 32 kernels
96f71476
Add env variable ORT_DISABLE_FUSED_ATTENTION
3e7f3f81
exclude files in hipify
0416aac6
fix cpplint; AttentionPastState_dynamic threshold
f327fec0
format
5aa9e2e2
fix --use_mask_index in benchmark
b1440e56
Merge branch 'main' into tlwu/trt_fused_attention
52ff22e1
tianleiwu
marked this pull request as ready for review 3 years ago
gh-yewang
approved these changes
on 2022-09-13
tianleiwu
merged
95c4fc68
into main 3 years ago
tianleiwu
deleted the tlwu/trt_fused_attention branch 3 years ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub