transformers
7d2a05a8 - [generate] Fix CUDAGraphs tensor overwrite for encoder-decoder static cache

Commit
49 days ago
[generate] Fix CUDAGraphs tensor overwrite for encoder-decoder static cache `_prepare_static_cache` now calls `early_initialization` on BOTH the self-attention and cross-attention caches for encoder-decoder models. Previously only the self-attention cache was initialized eagerly. For encoder-decoder models (e.g. T5), cross-attention K/V tensors are computed from encoder hidden states on the first decode step, which happens inside the CUDAGraph capture context. When `lazy_initialization` runs there, `is_torchdynamo_compiling()` is True so `mark_static_address` is skipped. CUDAGraphs then pools the tensor memory addresses, causing: RuntimeError: Error: accessing tensor output of CUDAGraphs that has been overwritten by a subsequent run. Regression introduced by #47731 ("Stop setting the static cache as an attribute to save memory"), which changed `_prepare_static_cache` to always create a fresh `StaticCache` instead of reusing `self._cache`. Before that PR, the same cache object was reused across `generate` calls and was already initialized from the first call (outside any compiled context), so `lazy_initialization` never ran inside capture. Also folds the decoder-only early-init (previously only for chunked prefill, #46421) into the same unconditional block, since the same CUDAGraph issue applies to any compiled generate call. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Author
Committer
Parents
Loading