onnxruntime
[CUDA] PagedAttention: add SM<80 fp16 fallback via memory-efficient attention
#28200
Merged

Loading