Fix GPU memory teardown in CLI serve tests (#48618)
* [fix] Free GPU memory between CLI test classes to prevent meta-device OOM
Each tearDownClass now calls cleanup() + sets cls.serve = None to release
GPU memory before the next test class loads a model. Also enables
expandable_segments to reduce fragmentation from sequential large model loads.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* [fix] Set expandable_segments via os.environ before CUDA init in CLI tests
set_allocator_settings() is a no-op once the CUDA allocator is first used,
which happens at import time. Setting PYTORCH_CUDA_ALLOC_CONF via os.environ
at module load time (conftest.py top-level) ensures it takes effect before
any CUDA allocation occurs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* [fix] Release InferenceThread closure references before blocking on queue.get()
While InferenceThread blocks on queue.get() waiting for the next task, the
local variable `fn` from the previous iteration still holds the submitted
closure. For inference calls this closure captures `model` via
`model.generate(**gen_kwargs)`, keeping the model's GPU tensors alive even
after delete_model() clears all other references.
Explicitly set fn/args/kwargs = None in a finally block so the closure (and
the model it captures) is released before the thread blocks again.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* [fix] Clean up: remove unneeded expandable_segments env var, add cleanup comment
The PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True was added to work around
fragmentation, but the InferenceThread fix fully resolves the root cause so it
is not needed. Also adds a short comment on cleanup() calls in tearDownClass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>