Document and test packed token inputs for MoE and QMoE (#32522)
## Description
Clarify and test packed token-major input support for the `MoE` and
`QMoE`
operators on the CPU and CUDA execution providers.
Packed inputs use:
- Input: `[total_tokens, hidden_size]`
- Router probabilities: `[total_tokens, num_experts]`
- Output: `[total_tokens, hidden_size]`
Tokens from variable-length sequences can be concatenated without
padding
because MoE routing is token-local and does not require sequence
boundaries.
## Changes
- Update the `MoE` and `QMoE` schema documentation.
- Document packed input behavior for CPU and CUDA.
- Add CPU parity tests for packed `MoE` and INT4 `QMoE`.
- Add CUDA parity tests for packed `MoE` and INT4 `QMoE`.
- Fix the CPU test harness to pass processed routing probabilities to
`MoE`
while preserving raw router logits for `QMoE`.
No runtime kernel changes are required because the existing CPU and CUDA
implementations already process rank-2 token-major inputs.