[CUDA] Skip FP4 QMoE fc1 activation expansion (#31479)
## Description
Stacked on #31159; review only the top commit.
This removes the standalone FP4 QMoE fc1 activation expansion during
GEMV decode. Instead, fc1 maps each permuted row back to its source
token with `permuted_row_to_source_row[row] % num_rows`; fc2 remains
unchanged because it consumes the expanded fc1 output.
## Summary of Changes
- Pass the permuted-row-to-source-row mapping through all FP4
interleaved SwiGLU GEMV launch variants.
- Read fc1 activations directly from the original input while preserving
PR 31159's SM80 pair-interleaved weight layout.
- Keep the legacy path available with
`ORT_DISABLE_FP4_GEMV_SKIP_EXPAND=1` for same-binary comparison.
- Add an exact-output parity test for the MTP shape (`num_tokens=3`,
`top_k=8`) with skip-expand enabled and disabled.
## Performance
H200, Qwen3.6 35B A3B NVFP4, MTP N=3, paired same-binary A/B:
- Removes 40 `expandInputRows` launches per decoding step.
- Removes 94.651 us/step of activation expansion.
- Adds 10.377 us/step to fc1 source-row lookup.
- Saves 84.274 us/step across named kernels.
- Saves 0.097695 ms/step median end-to-end (1.314%).
## Testing
- `lintrunner -a` on the five changed files.
- Compiled `moe_gemv_fp4.cu` and `moe_quantization.cc` against the exact
#31159 head.
- New expanded-vs-skip-expand regression passes with exact tensor
equality.
- 23 NVFP4 QMoE CUDA tests pass locally; the separately isolated
large-input scaling test fails identically with skip-expand enabled and
disabled on the pre-existing integration binary.
## Checklist
- [x] Tests added/updated
- [x] No breaking changes
- [x] No documentation changes required