Fix Dynamo Benchmark: Eager Deepcopy Inflating Memory Compression (#196178)
Summary:
## Summary
Correct Dynamo benchmark eager-memory accounting by ensuring the eager and
compiled measurements have the same number of resident model copies.
## Bug
The performance runner created the compiled model copy before eager warmup.
Eager warmup then ran on another deepcopy(model) to prevent stateful
operations such as generate() from mutating the model used for compilation.
Consequently, memory measurement included:
- Eager: loaded model + compiled copy + eager copy
- Compiled: loaded model + compiled copy
This extra weight copy inflated eager_peak_mem and therefore
compression_ratio.
## Fix
Run eager warmup using an isolated model before creating the compiled model:
1. Create the eager model and its optimizer.
2. Perform the existing eager warmups and memory measurement.
3. Release the eager optimizer/model and clear the device cache.
4. Create a fresh compiled model from the unmodified loaded model.
This preserves model-state isolation for generate() and other stateful
workloads. It also preserves the existing measured/stabilization iteration
accounting and adds no forward passes.
The change applies to models using the shared Dynamo performance runner,
including TorchBench, timm, and HuggingFace models.
## Memory impact
Qwen3-0.6B, BF16, 1,000 input tokens and 2,000 generated tokens:
| Metric | Before | After | Change |
|---|---|---|---|
| Eager peak memory | 3.959 GB | 2.765 GB | -1.194 GB |
| Compiled peak memory | 2.385 GB | 2.385 GB | unchanged |
| Compression ratio | 1.659x | 1.159x | -0.500x |
The 1.194 GB reduction closely matches the model’s 1.192 GB of BF16
parameters, confirming that the extra weight copy was removed.
## Test plan
python -m unittest \
benchmarks.dynamo.test.TestDynamoBenchmark.test_eager_warmup_does_not_retain_c
ompiled_model \
benchmarks.dynamo.test.TestHuggingFaceLLMPerformance.test_compilation_latency_
uses_matched_work
python -m py_compile benchmarks/dynamo/common.py benchmarks/dynamo/test.py
Also measured Qwen3-0.6B before and after using dashboard-style warm peak-
memory collection.
X-link: https://github.com/pytorch/pytorch/pull/196178
Approved by: https://github.com/anijain2305
Reviewed By: izaitsevfb
Differential Revision: D119256599
fbshipit-source-id: 6b0d2d24c005da1dc64a2c33ddd00ac2603f7c93