transformers
0a959de1 - Fix GPU memory teardown in CLI serve tests (#48618)

Commit
29 days ago
Fix GPU memory teardown in CLI serve tests (#48618) * [fix] Free GPU memory between CLI test classes to prevent meta-device OOM Each tearDownClass now calls cleanup() + sets cls.serve = None to release GPU memory before the next test class loads a model. Also enables expandable_segments to reduce fragmentation from sequential large model loads. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [fix] Set expandable_segments via os.environ before CUDA init in CLI tests set_allocator_settings() is a no-op once the CUDA allocator is first used, which happens at import time. Setting PYTORCH_CUDA_ALLOC_CONF via os.environ at module load time (conftest.py top-level) ensures it takes effect before any CUDA allocation occurs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [fix] Release InferenceThread closure references before blocking on queue.get() While InferenceThread blocks on queue.get() waiting for the next task, the local variable `fn` from the previous iteration still holds the submitted closure. For inference calls this closure captures `model` via `model.generate(**gen_kwargs)`, keeping the model's GPU tensors alive even after delete_model() clears all other references. Explicitly set fn/args/kwargs = None in a finally block so the closure (and the model it captures) is released before the thread blocks again. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [fix] Clean up: remove unneeded expandable_segments env var, add cleanup comment The PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True was added to work around fragmentation, but the InferenceThread fix fully resolves the root cause so it is not needed. Also adds a short comment on cleanup() calls in tearDownClass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Author
Parents
Loading