LifeStream: Token Life-Cycle Modeling for Training-Free Online Video Understanding
Abstract
With the rapid development of Video Large Language Models (VLLMs), online video understanding has emerged as a practical paradigm for real-world streaming applications. A key challenge is to construct a compact, query-agnostic visual memory from continuously growing video streams under strict token budgets. Existing token reduction methods typically rely on static and coarse-grained importance estimation, implicitly assuming that retained tokens contribute equally over time. This overlooks the dynamic nature of visual information, where the relevance of tokens evolves as the video progresses, leading to suboptimal prioritization and inefficient long-term memory usage. To address this limitation, we propose LifeStream, a training-free framework for lifecycle-aware visual token selection and hierarchical memory construction. LifeStream models the temporal evolution of visual tokens and reveals that tokens contribute unequally to downstream reasoning depending on their lifecycle states. Based on this insight, it organizes tokens into a hierarchical memory structure: informative non-stable tokens are preserved in an archive memory, while long-range history is progressively compressed into a global memory according to state-specific priorities. This design enables continuous retention of high-value information under limited token budgets, without requiring query-time retrieval or additional computation. Extensive experiments demonstrate that LifeStream discards over 80\% of visual tokens while improving performance. It improves LLaVA-OV-7B and Qwen3-VL-8B by 4.4\% and 3.5\% on StreamingBench, respectively, and achieves state-of-the-art results on one additional online benchmark and four long-video understanding benchmarks.