[CUDA] Give each plugin EP session its own device arena (#32807)
### Description
The CUDA Plugin EP factory handed **one BFC device arena per GPU** to
every `CreateAllocator` caller: every session, plus the environment's
shared allocator. After a CUDA graph is captured, the session frees the
graph's intermediate chunks back to the arena, but the captured graph
still reads and writes those addresses on every replay. Because the
arena was shared, another session could be given those chunks and
overwrite them between replays.
With two graph-captured sessions on one device (a Qwen3.8-27B target
plus a DFlash2 drafter in onnxruntime-genai), this silently corrupted
generation. MMLU-Pro dropped 5.5 pp and speculative acceptance fell from
87.7% to 69.8%. The bundled CUDA EP is not affected because it creates a
separate arena for each session.
This PR gives each `CreateAllocator` call for device memory its own
arena, which matches the bundled EP. It replaces the temporary
graph-disable gate in #32801.
#### Changes
| File | Change |
|------|--------|
| `cuda_ep_factory.h/.cc` | Replace the single `device_arena` and its
refcount with `std::vector<DeviceArena> device_arenas`. Each entry
tracks its own `has_quarantine` / `abandoned` state.
`CreateAllocatorImpl` creates a new arena on every call and now honors
that session's arena options, which the shared arena used to ignore with
a warning. `ReleaseAllocatorImpl` destroys only the matching arena, or
leaks it on purpose if it is quarantined or abandoned. Stream run-end
reset and quarantine/abandon after an undrained release now visit every
arena on the device. For abandon, this is the same scope the single
shared arena had before. `GetDeviceArenaForDevice()` had no callers and
is removed. |
| `cuda_plugin_arena_test.cc` | New `DeviceAllocator_IsSessionScoped`
test: an allocation from session A must not show up in session B's arena
stats or the environment allocator's stats. It **fails on the current
shared-arena plugin** (`NumAllocs` 3 vs 2) and passes with this change.
|
| `docs/cuda_plugin_ep/*.md` | Describe the per-session device arena and
why CUDA graphs require it. |
The pinned arena and the CUDA mempool allocator are still shared per
device and are unchanged.
### Motivation and Context
#### Root-cause isolation (MMLU-Pro, 200 questions, 8,192 max new
tokens, DFlash2 drafts=6, concurrency 8)
| Configuration | Correct | Cap hits | Acceptance |
|---|---:|---:|---:|
| Bundled CUDA EP, graphs on | 168 | 3 | 87.7% |
| Plugin (current), graphs on | 157 | 64 | 69.8% |
| Plugin, target session graph only | 156 | 66 | 69.4% |
| Plugin, drafter session graph only | 167 | 7 | 86.2% |
| Plugin, per-session arena prototype, both graphs on | 169 | 4 | 87.8%
|
| Same prototype binary, shared arena (control) | 156 | 64 | 70.1% |
Toggling only the arena sharing, with the same binary and both graphs
on, flips the result. That identifies the shared arena as the cause.
#### This PR (graphs on for both sessions)
| Metric | Bundled | Plugin with this PR |
|---|---:|---:|
| MMLU-Pro correct | 168/200 | **168/200** |
| Cap hits | 3 | 3 |
| Acceptance | 87.7% | 87.8% |
| Paired vs bundled | — | 1 win / 1 loss, McNemar p=1.0, 183/200
token-exact |
| Paired vs current plugin | — | 13 wins / 2 losses, p=0.0074 |
#### Performance
Batch 1, 32,768 prompt tokens, 1,024 generated tokens, DFlash2 draft
width 6, 3 fresh-process repetitions, all arms run at the same time on
separate H200s:
| Runtime | Median decode tok/s | Median prefill tok/s | Median TTFT |
Peak workload |
|---|---:|---:|---:|---:|
| Bundled CUDA EP | 314.58 | 3,767.2 | 8.698 s | 31,590 MiB |
| Plugin, graphs disabled (#32801 gate) | 305.65 (0.972x) | 3,822.9 |
8.572 s | 30,900 MiB |
| **Plugin, this PR** | **317.01 (1.008x)** | 3,780.7 | 8.667 s | 31,526
MiB |
Re-enabling graphs safely recovers the 2.8% decode loss from the gate.
Peak memory is 64 MiB below the bundled EP.
#### Tests
- `CudaPluginArenaTest.*`, `CudaPluginUserStreamGraphTest.*`,
`CudaPluginDeviceDiscoveryTest.*`: 51/51 passed on an H200 (Release,
SM90).
- `clang-format --dry-run --Werror` and `git diff --check` pass.