Update MoE offload implementation plan and implements a global state for MoE/QMoE (#32738)
## Description
- update the MoE CPU-offload implementation plan and document counter
configuration and the initial-state format
- add opt-in session-global expert state for CPU and built-in CUDA
MoE/QMoE, shared with subgraphs and isolated between sessions
- build the immutable kernel/expert index at initialization; keep global
counters, reads, and updates in MoeExpertState
- own one concrete KernelUsage collector per registered kernel in the
session state; CPU and CUDA obtain it through
OpKernelContext::GetKernelUsage and collect expert IDs directly
- forward only the collector getter through the internal shared-provider
bridge, without extending the public C API; keep copies, pinned buffers,
and events in CudaRoutingSnapshot
- commit collected usage in the executor after successful kernel
execution; failed kernels do not update counters, and invocations that
do not access collection do not replay stale selections
- update counters in place once per node invocation; require finite
non-negative alpha and beta with alpha + beta <= 1, defaulting to 0.9
and 0.1
- use no RecordUsage mutex and reject overlapping runs on the same
session
- overlap CUDA counting with expert computation: enqueue a reusable
pinned routing snapshot and completion event after top-k, then collect
usage on the CPU after launching expert kernels, waiting only for the
snapshot rather than the whole stream
- preserve fused QMoE routing and union selected experts across row
tiles; cover buffer reuse and aborted-invocation recovery
- exclude counting state and collectors from minimal builds; explicitly
reject unsupported configurations, including the CUDA plugin EP
This PR implements the session-global state and counting stage. It does
not enable expert placement, CPU offload, swaps, or budget
redistribution.
## Validation
- CPU and built-in CUDA targets build successfully.
- 35 focused framework/collector/state/counting tests passed on H200,
including context collector ownership, failed-kernel commit rejection
and all 4 CUDA integration cases with the default routing configuration.
- All 4 CUDA counting integration cases also passed with fused routing
disabled.
- 6 real-CUDA collector tests passed, covering nonblocking collection,
pending device work, tiled union, buffer reuse, aborted-call cleanup,
and invalid/uninitialized input handling.
- Generic collector tests cover deduplication, independent instances,
atomic invalid-batch rejection, reset, and retained storage across
invocations.
- 3 native minimal-build tests passed.
- C++ formatting and git diff --check passed.
- The earlier minimal Android size fix measured 1,594,893 bytes against
the unchanged 1,595,392-byte threshold.
- Rename-only follow-up: verified identifier/path-only changes and
syntax/type-checked CUDA MoE, QMoE, and the renamed routing snapshot
tests with the configured compiler flags.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>