onnxruntime
d5a52384 - Document and test packed token inputs for MoE and QMoE (#32522)

Commit
14 days ago
Document and test packed token inputs for MoE and QMoE (#32522) ## Description Clarify and test packed token-major input support for the `MoE` and `QMoE` operators on the CPU and CUDA execution providers. Packed inputs use: - Input: `[total_tokens, hidden_size]` - Router probabilities: `[total_tokens, num_experts]` - Output: `[total_tokens, hidden_size]` Tokens from variable-length sequences can be concatenated without padding because MoE routing is token-local and does not require sequence boundaries. ## Changes - Update the `MoE` and `QMoE` schema documentation. - Document packed input behavior for CPU and CUDA. - Add CPU parity tests for packed `MoE` and INT4 `QMoE`. - Add CUDA parity tests for packed `MoE` and INT4 `QMoE`. - Fix the CPU test harness to pass processed routing probabilities to `MoE` while preserving raw router logits for `QMoE`. No runtime kernel changes are required because the existing CPU and CUDA implementations already process rank-2 token-major inputs.
Author
Parents
Loading