[CUDA] Fix QMoE int4/int8 weight prepack to always use SM80 layout (#28978)
## Summary
The CUDA QMoE INT4/INT8 grouped GEMM always dispatches to the Ampere
(SM80) CUTLASS kernel — even on Hopper (SM90) — because mixed int-weight
+ fp16/bf16 activation is not a valid Hopper TMA warp-specialized
specialisation. This PR makes weight prepacking always emit the SM80
(column-interleaved) `fpA_intB` layout regardless of the runtime device
SM, fixing silently-wrong output on Hopper, and centralizes the
arch-clamping logic in a single shared helper. It also cleans up the
related tests and tightens MoE parity tolerances that were too loose to
catch the layout bug.
## Motivation
https://github.com/microsoft/onnxruntime/pull/28749 uses 90 for sm90
weight prepacking.
On SM90, `isValidHopperMOESpecialisation<half_t, uint4b_t/uint8_t>()` is
`false`, so the grouped MoE GEMM falls back to the SM80 kernel. The
weight preprocessor, however, skips column interleaving for `arch ==
90`, so an auto-detected (`force_arch=-1`) pack on an H200 produced the
non-interleaved SM90 layout that the SM80 kernel cannot consume —
yielding wrong results. The previous `PrePackIntExpertWeights` logic
clamped to `sm_` (passing SM90 through), and the test that exercised the
offline packer used auto-detect, so both could emit the wrong layout.
## Key Changes
| Area | Change |
|---|---|
| `fpA_intB_gemm_preprocessors{.h,_impl.cu}` | Extracted
`get_arch_for_mixed_gemm_weight_preprocess(int arch)` as a shared,
declared helper (clamps SM to the layout group: `<80→75`, `90→90`, else
`80`). |
| `fpA_intB_gemm_preprocessors_impl.h` | `getLayoutDetailsForTransform`
now routes through the shared helper instead of duplicating the
arch-range logic. |
| `moe_quantization.cc` (`PrePackIntExpertWeights`) | Always packs
INT4/INT8 expert weights for the SM80 layout
(`get_arch_for_mixed_gemm_weight_preprocess(80)`) instead of clamping to
the runtime `sm_`, since the SM80 kernel runs on every GPU. |
| `onnxruntime_pybind_quant.cc` (`PackWeightsForMixedGemm`) | Replaced
the ad-hoc `{75,80,90}` allowlist with the shared helper, so
`force_arch` is clamped consistently with the runtime dispatch (removes
the now-unused `<set>` include). |
| `contrib_defs.cc` / `moe_quantization.h` | Updated `weights_prepacked`
schema/field docs: layouts for `-1`/`1` are EP-determined; for the CUDA
EP `-1` and `1` are equivalent today (both SM80), `1` reserved for a
future Hopper-specific layout. |
| `test_qmoe_cuda.py` | Removed the dead, never-called
`preprocess_weights_for_mixed_gemm` helper; the real path
(`quant_dequant_blockwise`) already pins `sm=80`. |
| `test_moe_cuda.py` | Pinned the offline packer to `arch=80`, and
tightened FP16 QMoE parity tolerance from `atol 3.0 (4-bit)` / `2.0
(8-bit)` to `0.5` now that the layout is correct. |
| `docs/` | Regenerated `ContribOperators.md` and updated `moe_qmoe.md`
to match the new schema docs and SM80-always packing rationale. |
## Testing Notes
On an H200 (SM90), with the CUDA 12.x/13.x Python wheel:
```bash
python -m pytest onnxruntime/test/python/transformers/test_qmoe_cuda.py
python -m pytest onnxruntime/test/python/transformers/test_moe_cuda.py -k "PhiQMoE or qmoe"
```
- `test_qmoe_cuda.py` SwiGLU parity: SM80 layout → max diff ~0.001
(pass, tol 0.1); the prior SM90 layout produced max diff ~1.2 (fail),
confirming the fix.
- `test_moe_cuda.py` `TestPhiQMoE` (4-bit and 8-bit, all batch/seq
combinations): worst observed `max_diff` ≈ 0.375 with the fixed layout,
comfortably under the new `atol=0.5`.
- `ruff check` passes on both edited test files.
---------
Co-authored-by: tlwu <tlwu@example.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>