Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
Abstract
Online video understanding requires models to robustly encode, store, and retrieve information from a continuous stream to support accurate question answering. Existing approaches rely on key--value caching to accumulate frame-level representations over time, but allocate only a limited token budget per frame, causing loss of fine-grained visual detail. We find that naively scaling this budget degrades retrieval quality and question-answering performance, as denser token streams introduce redundancy that disrupts query--frame similarity. To address this, we first introduce an adaptive selection strategy that reduces token redundancy while preserving local spatiotemporal information. Second, we propose a training-free retrieval ensemble that leverages external models to better identify relevant frames. Our method, \textbf{MemStream}, achieves +8.0\% on CG-Bench, +8.5\% on LVBench, and +2.4\% on VideoMME (Long) over ReKV with Qwen2.5-VL-7B. Furthermore, we show that MemStream consistently improves performance across several Video-LLMs for long-video understanding.