onnxruntime
41426bd5 - Add H128 INT8 paged speculative XQA (#32740)

Commit
12 days ago
Add H128 INT8 paged speculative XQA (#32740) ## Summary This change adds a paged XQA specialization for speculative verification with: - FP16 query and output. - INT8 KV cache. - Head size 128. - GQA group size 6. - Query widths from 2 through 8. These shapes previously used a much slower portable multi-token attention backend. ## What changed - Add the H128 FP16-query/INT8-cache speculative XQA CUDA specialization. - Pass the head size through XQA launch validation, workspace sizing, and dynamic shared-memory sizing. - Route only the supported H128/group-six/INT8 shapes to the new specialization. - Keep the existing H256 speculative XQA paths unchanged. - Require causal attention because the speculative mask is lower triangular. - Keep unsupported widths, non-causal attention, and other unsupported shapes on the existing portable fallback. - Add tests for: - Query widths 2 through 8. - Ragged batches and uneven sequence lengths. - Fragmented page tables. - Local attention windows. - CUDA graph replay. - Width-9 fallback. - Non-causal fallback. - Existing H256 FP16, BF16, INT8, and FP8 paths. ## Performance At 128K context on the target SM121 GPU: | Path | Verification time | |---|---:| | Portable fallback | 13.7-14.0 ms | | H128 INT8 speculative XQA | 1.13-1.14 ms | This is about a **12x reduction in attention verification latency** for supported widths 2-8. This result measures the attention operator. End-to-end speculative decoding speedup also depends on draft acceptance, proposal cost, and the rest of the model. --------- Copilot-Session: 0973b598-cc85-4145-93d7-e4459d07f7b7
Author
Parents
Loading