onnxruntime
7843226c - [CUDA] Skip FP4 QMoE fc1 activation expansion (#31479)

Commit
51 days ago
[CUDA] Skip FP4 QMoE fc1 activation expansion (#31479) ## Description Stacked on #31159; review only the top commit. This removes the standalone FP4 QMoE fc1 activation expansion during GEMV decode. Instead, fc1 maps each permuted row back to its source token with `permuted_row_to_source_row[row] % num_rows`; fc2 remains unchanged because it consumes the expanded fc1 output. ## Summary of Changes - Pass the permuted-row-to-source-row mapping through all FP4 interleaved SwiGLU GEMV launch variants. - Read fc1 activations directly from the original input while preserving PR 31159's SM80 pair-interleaved weight layout. - Keep the legacy path available with `ORT_DISABLE_FP4_GEMV_SKIP_EXPAND=1` for same-binary comparison. - Add an exact-output parity test for the MTP shape (`num_tokens=3`, `top_k=8`) with skip-expand enabled and disabled. ## Performance H200, Qwen3.6 35B A3B NVFP4, MTP N=3, paired same-binary A/B: - Removes 40 `expandInputRows` launches per decoding step. - Removes 94.651 us/step of activation expansion. - Adds 10.377 us/step to fc1 source-row lookup. - Saves 84.274 us/step across named kernels. - Saves 0.097695 ms/step median end-to-end (1.314%). ## Testing - `lintrunner -a` on the five changed files. - Compiled `moe_gemv_fp4.cu` and `moe_quantization.cc` against the exact #31159 head. - New expanded-vs-skip-expand regression passes with exact tensor equality. - 23 NVFP4 QMoE CUDA tests pass locally; the separately isolated large-input scaling test fails identically with skip-expand enabled and disabled on the pre-existing integration binary. ## Checklist - [x] Tests added/updated - [x] No breaking changes - [x] No documentation changes required
Author
Parents
Loading