[generate] Fix CUDAGraphs tensor overwrite for encoder-decoder static cache
`_prepare_static_cache` now calls `early_initialization` on BOTH the
self-attention and cross-attention caches for encoder-decoder models.
Previously only the self-attention cache was initialized eagerly. For
encoder-decoder models (e.g. T5), cross-attention K/V tensors are
computed from encoder hidden states on the first decode step, which
happens inside the CUDAGraph capture context. When `lazy_initialization`
runs there, `is_torchdynamo_compiling()` is True so `mark_static_address`
is skipped. CUDAGraphs then pools the tensor memory addresses, causing:
RuntimeError: Error: accessing tensor output of CUDAGraphs that has
been overwritten by a subsequent run.
Regression introduced by #47731 ("Stop setting the static cache as an
attribute to save memory"), which changed `_prepare_static_cache` to
always create a fresh `StaticCache` instead of reusing `self._cache`.
Before that PR, the same cache object was reused across `generate`
calls and was already initialized from the first call (outside any
compiled context), so `lazy_initialization` never ran inside capture.
Also folds the decoder-only early-init (previously only for chunked
prefill, #46421) into the same unconditional block, since the same
CUDAGraph issue applies to any compiled generate call.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>