onnxruntime
4c868628 - [CUDA] Give each plugin EP session its own device arena (#32807)

Commit
14 days ago
[CUDA] Give each plugin EP session its own device arena (#32807) ### Description The CUDA Plugin EP factory handed **one BFC device arena per GPU** to every `CreateAllocator` caller: every session, plus the environment's shared allocator. After a CUDA graph is captured, the session frees the graph's intermediate chunks back to the arena, but the captured graph still reads and writes those addresses on every replay. Because the arena was shared, another session could be given those chunks and overwrite them between replays. With two graph-captured sessions on one device (a Qwen3.8-27B target plus a DFlash2 drafter in onnxruntime-genai), this silently corrupted generation. MMLU-Pro dropped 5.5 pp and speculative acceptance fell from 87.7% to 69.8%. The bundled CUDA EP is not affected because it creates a separate arena for each session. This PR gives each `CreateAllocator` call for device memory its own arena, which matches the bundled EP. It replaces the temporary graph-disable gate in #32801. #### Changes | File | Change | |------|--------| | `cuda_ep_factory.h/.cc` | Replace the single `device_arena` and its refcount with `std::vector<DeviceArena> device_arenas`. Each entry tracks its own `has_quarantine` / `abandoned` state. `CreateAllocatorImpl` creates a new arena on every call and now honors that session's arena options, which the shared arena used to ignore with a warning. `ReleaseAllocatorImpl` destroys only the matching arena, or leaks it on purpose if it is quarantined or abandoned. Stream run-end reset and quarantine/abandon after an undrained release now visit every arena on the device. For abandon, this is the same scope the single shared arena had before. `GetDeviceArenaForDevice()` had no callers and is removed. | | `cuda_plugin_arena_test.cc` | New `DeviceAllocator_IsSessionScoped` test: an allocation from session A must not show up in session B's arena stats or the environment allocator's stats. It **fails on the current shared-arena plugin** (`NumAllocs` 3 vs 2) and passes with this change. | | `docs/cuda_plugin_ep/*.md` | Describe the per-session device arena and why CUDA graphs require it. | The pinned arena and the CUDA mempool allocator are still shared per device and are unchanged. ### Motivation and Context #### Root-cause isolation (MMLU-Pro, 200 questions, 8,192 max new tokens, DFlash2 drafts=6, concurrency 8) | Configuration | Correct | Cap hits | Acceptance | |---|---:|---:|---:| | Bundled CUDA EP, graphs on | 168 | 3 | 87.7% | | Plugin (current), graphs on | 157 | 64 | 69.8% | | Plugin, target session graph only | 156 | 66 | 69.4% | | Plugin, drafter session graph only | 167 | 7 | 86.2% | | Plugin, per-session arena prototype, both graphs on | 169 | 4 | 87.8% | | Same prototype binary, shared arena (control) | 156 | 64 | 70.1% | Toggling only the arena sharing, with the same binary and both graphs on, flips the result. That identifies the shared arena as the cause. #### This PR (graphs on for both sessions) | Metric | Bundled | Plugin with this PR | |---|---:|---:| | MMLU-Pro correct | 168/200 | **168/200** | | Cap hits | 3 | 3 | | Acceptance | 87.7% | 87.8% | | Paired vs bundled | — | 1 win / 1 loss, McNemar p=1.0, 183/200 token-exact | | Paired vs current plugin | — | 13 wins / 2 losses, p=0.0074 | #### Performance Batch 1, 32,768 prompt tokens, 1,024 generated tokens, DFlash2 draft width 6, 3 fresh-process repetitions, all arms run at the same time on separate H200s: | Runtime | Median decode tok/s | Median prefill tok/s | Median TTFT | Peak workload | |---|---:|---:|---:|---:| | Bundled CUDA EP | 314.58 | 3,767.2 | 8.698 s | 31,590 MiB | | Plugin, graphs disabled (#32801 gate) | 305.65 (0.972x) | 3,822.9 | 8.572 s | 30,900 MiB | | **Plugin, this PR** | **317.01 (1.008x)** | 3,780.7 | 8.667 s | 31,526 MiB | Re-enabling graphs safely recovers the 2.8% decode loss from the gate. Peak memory is 64 MiB below the bundled EP. #### Tests - `CudaPluginArenaTest.*`, `CudaPluginUserStreamGraphTest.*`, `CudaPluginDeviceDiscoveryTest.*`: 51/51 passed on an H200 (Release, SM90). - `clang-format --dry-run --Werror` and `git diff --check` pass.
Author
Parents
Loading