Add Paged Attention Op for CUDA SM80 support #24595
paged attention op
6e75e228
test file amid-debug
aa6fd44c
paged attention works
5583aaaf
everything works and is implemented
c628e66b
small stuff
554d0c70
lint
497a22eb
address comments
c3276d11
fix dmmha rotary and address tianleiwu comments
482240c8
increase efficiency
6a197dc6
correction
0f23e6f8
update flash and remove redundant inputs
cd15f813
bert comment
aa9b6f19
merge main
e0432f8d
docs
50a3767d
comments
5eccad2c
docs
1c3b53c2
tianleiwu
approved these changes
on 2025-06-12
aciddelgado
deleted the aciddelgado/paged_attention branch 1 year ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub