[CUDA] Add FP16-cache paged XQA decode for head size 256 #32263
baijumeswani
force pushed
from
75ba73ae
to
af085e8e
30 days ago
Optimize FP16 long-context paged decode with XQA
7c592b8b
baijumeswani
force pushed
from
af085e8e
to
7c592b8b
30 days ago
tianleiwu
approved these changes
on 2026-08-26
baijumeswani
deleted the baijumeswani/fp16-paged-xqa-pr branch 29 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub