Add support for grouped-query MultiHeadAttention for CPU and CUDA #31680
tianleiwu
changed the title Add CPU support for grouped-query MultiHeadAttention Add support for grouped-query MultiHeadAttention for CPU and CUDA 60 days ago
tianleiwu
force pushed
from
cddb69e2
to
785368c2
52 days ago
tianleiwu
marked this pull request as ready for review 46 days ago
Add grouped-query MHA validation and CUDA path
fc7c88d8
Complete grouped-query MHA shapes and cache handling
d09280db
Finalize grouped-query MHA validation and tests
51dfb9e6
Implement CPU grouped-query attention
e5c0758a
Fix grouped-query shape inference build
88111b7f
address feedbacks
2b43d905
Address review feedback on grouped-query MultiHeadAttention
7688cfe8
Pass kv_num_heads explicitly to ApplyAttention
6da01191
Add hidden size accessors to AttentionParameters
c83a7809
Replace AttentionParameters::v_hidden_size with hidden size accessors
cbec80dd
fix build
887b17f7
consistent naming
54ba92b7
fix(webgpu): use grouped-query value dimensions
db4e5734
docs: regenerate contrib operator reference
aff244cc
docs: fix grouped-query attribute placement
555fb892
fix(attention): correct grouped-query caches
d1bd6a98
tianleiwu
force pushed
from
c6831587
to
d1bd6a98
33 days ago
Keep allowing bias when key and value are 4-D BNSH tensors
d574822b
tianleiwu
force pushed
from
aa50ec59
to
d574822b
31 days ago
fix(attention): address GQA review feedback
83144876
tianleiwu
dismissed their stale review
via 83144876
30 days ago
fix(attention): validate GQA head counts in symbolic shape inference
ced57815
Login to write a write a comment.
Login via GitHub