onnxruntime
c2650ccb - Update MoE offload implementation plan and implements a global state for MoE/QMoE (#32738)

Commit
7 days ago
Update MoE offload implementation plan and implements a global state for MoE/QMoE (#32738) ## Description - update the MoE CPU-offload implementation plan and document counter configuration and the initial-state format - add opt-in session-global expert state for CPU and built-in CUDA MoE/QMoE, shared with subgraphs and isolated between sessions - build the immutable kernel/expert index at initialization; keep global counters, reads, and updates in MoeExpertState - own one concrete KernelUsage collector per registered kernel in the session state; CPU and CUDA obtain it through OpKernelContext::GetKernelUsage and collect expert IDs directly - forward only the collector getter through the internal shared-provider bridge, without extending the public C API; keep copies, pinned buffers, and events in CudaRoutingSnapshot - commit collected usage in the executor after successful kernel execution; failed kernels do not update counters, and invocations that do not access collection do not replay stale selections - update counters in place once per node invocation; require finite non-negative alpha and beta with alpha + beta <= 1, defaulting to 0.9 and 0.1 - use no RecordUsage mutex and reject overlapping runs on the same session - overlap CUDA counting with expert computation: enqueue a reusable pinned routing snapshot and completion event after top-k, then collect usage on the CPU after launching expert kernels, waiting only for the snapshot rather than the whole stream - preserve fused QMoE routing and union selected experts across row tiles; cover buffer reuse and aborted-invocation recovery - exclude counting state and collectors from minimal builds; explicitly reject unsupported configurations, including the CUDA plugin EP This PR implements the session-global state and counting stage. It does not enable expert placement, CPU offload, swaps, or budget redistribution. ## Validation - CPU and built-in CUDA targets build successfully. - 35 focused framework/collector/state/counting tests passed on H200, including context collector ownership, failed-kernel commit rejection and all 4 CUDA integration cases with the default routing configuration. - All 4 CUDA counting integration cases also passed with fused routing disabled. - 6 real-CUDA collector tests passed, covering nonblocking collection, pending device work, tiled union, buffer reuse, aborted-call cleanup, and invalid/uninitialized input handling. - Generic collector tests cover deduplication, independent instances, atomic invalid-batch rejection, reset, and retained storage across invocations. - 3 native minimal-build tests passed. - C++ formatting and git diff --check passed. - The earlier minimal Android size fix measured 1,594,893 bytes against the unchanged 1,595,392-byte threshold. - Rename-only follow-up: verified identifier/path-only changes and syntax/type-checked CUDA MoE, QMoE, and the renamed routing snapshot tests with the configured compiler flags. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Author
Parents
Loading