Add H128 INT8 paged speculative XQA (#32740)
## Summary
This change adds a paged XQA specialization for speculative verification
with:
- FP16 query and output.
- INT8 KV cache.
- Head size 128.
- GQA group size 6.
- Query widths from 2 through 8.
These shapes previously used a much slower portable multi-token
attention backend.
## What changed
- Add the H128 FP16-query/INT8-cache speculative XQA CUDA
specialization.
- Pass the head size through XQA launch validation, workspace sizing,
and dynamic shared-memory
sizing.
- Route only the supported H128/group-six/INT8 shapes to the new
specialization.
- Keep the existing H256 speculative XQA paths unchanged.
- Require causal attention because the speculative mask is lower
triangular.
- Keep unsupported widths, non-causal attention, and other unsupported
shapes on the existing
portable fallback.
- Add tests for:
- Query widths 2 through 8.
- Ragged batches and uneven sequence lengths.
- Fragmented page tables.
- Local attention windows.
- CUDA graph replay.
- Width-9 fallback.
- Non-causal fallback.
- Existing H256 FP16, BF16, INT8, and FP8 paths.
## Performance
At 128K context on the target SM121 GPU:
| Path | Verification time |
|---|---:|
| Portable fallback | 13.7-14.0 ms |
| H128 INT8 speculative XQA | 1.13-1.14 ms |
This is about a **12x reduction in attention verification latency** for
supported widths 2-8.
This result measures the attention operator. End-to-end speculative
decoding speedup also depends
on draft acceptance, proposal cost, and the rest of the model.
---------
Copilot-Session: 0973b598-cc85-4145-93d7-e4459d07f7b7