benchmark
c27b8de9 - Fix Dynamo Benchmark: Eager Deepcopy Inflating Memory Compression (#196178)

Commit
19 days ago
Fix Dynamo Benchmark: Eager Deepcopy Inflating Memory Compression (#196178) Summary: ## Summary Correct Dynamo benchmark eager-memory accounting by ensuring the eager and compiled measurements have the same number of resident model copies. ## Bug The performance runner created the compiled model copy before eager warmup. Eager warmup then ran on another deepcopy(model) to prevent stateful operations such as generate() from mutating the model used for compilation. Consequently, memory measurement included: - Eager: loaded model + compiled copy + eager copy - Compiled: loaded model + compiled copy This extra weight copy inflated eager_peak_mem and therefore compression_ratio. ## Fix Run eager warmup using an isolated model before creating the compiled model: 1. Create the eager model and its optimizer. 2. Perform the existing eager warmups and memory measurement. 3. Release the eager optimizer/model and clear the device cache. 4. Create a fresh compiled model from the unmodified loaded model. This preserves model-state isolation for generate() and other stateful workloads. It also preserves the existing measured/stabilization iteration accounting and adds no forward passes. The change applies to models using the shared Dynamo performance runner, including TorchBench, timm, and HuggingFace models. ## Memory impact Qwen3-0.6B, BF16, 1,000 input tokens and 2,000 generated tokens: | Metric | Before | After | Change | |---|---|---|---| | Eager peak memory | 3.959 GB | 2.765 GB | -1.194 GB | | Compiled peak memory | 2.385 GB | 2.385 GB | unchanged | | Compression ratio | 1.659x | 1.159x | -0.500x | The 1.194 GB reduction closely matches the model’s 1.192 GB of BF16 parameters, confirming that the extra weight copy was removed. ## Test plan python -m unittest \ benchmarks.dynamo.test.TestDynamoBenchmark.test_eager_warmup_does_not_retain_c ompiled_model \ benchmarks.dynamo.test.TestHuggingFaceLLMPerformance.test_compilation_latency_ uses_matched_work python -m py_compile benchmarks/dynamo/common.py benchmarks/dynamo/test.py Also measured Qwen3-0.6B before and after using dashboard-style warm peak- memory collection. X-link: https://github.com/pytorch/pytorch/pull/196178 Approved by: https://github.com/anijain2305 Reviewed By: izaitsevfb Differential Revision: D119256599 fbshipit-source-id: 6b0d2d24c005da1dc64a2c33ddd00ac2603f7c93
Author
Committer
Parents
Loading