Learning-to-Memorize: Dynamic Context Management for Long-Horizon Autoregressive Video Generation
Abstract
Autoregressive video generation has achieved impressive results on short clips, yet generating long videos remains challenging due to error accumulation over extended horizons, where each predicted frame depends on potentially imperfect previous frames. Existing approaches typically rely on static context management strategies, such as sliding-window caches that discard early frames or fixed sink frames that permanently preserve initial content. These designs either lose critical long-range context or anchor generation to outdated information, which can lead to degraded fidelity, motion stagnation, and reduced diversity. In this paper, we propose \textbf{Learning to Memorize (L2M)}, a learned context memory management framework that dynamically retains, evicts, and evolves historical frames according to predicted importance. The framework consists of an \emph{importance prediction router} that estimates the relevance of each historical frame, \emph{degradation-aware training} that improves robustness to noisy or corrupted context, and a \emph{dynamic memory initialization and evolution} mechanism that updates long-term memory using exponential moving average (EMA) importance scores. This allows high-quality frames to replace stale anchors and enables memory to adapt naturally to scene evolution. Trained with a teacher-student distillation objective and sparsity regularization, L2M makes efficient use of a fixed KV-cache budget while preserving the most informative contextual information. Extensive experiments show that L2M achieves superior long-term video generation performance over existing baselines, improving stability and visual fidelity, and establishing a new paradigm for learned memory management in long autoregressive video generation.