MEMEVO: A Memory-Evolved Video Agent for Long Video Understanding
Abstract
Agent-based frameworks grounded in vision-language models (VLMs) have emerged as a dominant paradigm for long video understanding. Yet, prevailing agents lack the core capacity for \textit{dynamic memory evolution}, failing to transform ephemeral perceptions into a continuously growing, adaptive knowledge system and thus inducing inefficient redundant exploration during reasoning. To this end, we present MEMEVO, an online Memory-Evolved Video Agent framework centered on \textit{dynamic memory evolution} that leverages accumulating multi-turn queries to actively drive progressive memory evolution, establishing a long-lived memory system for knowledge accumulation. Central to MEMEVO is a four-level Hierarchical Self-Evolved Memory, which constructs memory bottom-up via progressive abstraction, distilling transient perceptions into reusable, query-agnostic event nodes organized within a temporal-semantic graph, while realizing active memory evolution through event insertion, duplicate merging, and redundant pruning. To mitigate redundant retrieval in the evolved memory, we devise a State-Conditioned Graph Memory Retrieval mechanism that guides efficient navigation of the memory space, integrating State-Conditioned Graph Routing and Meta-Cognitive Stopping criteria to ensure sufficient evidence acquisition without re-exploration. Extensive benchmarking confirms that MEMEVO achieves state-of-the-art performance on reasoning-intensive tasks (\textbf{\textcolor[HTML]{00FF00}{+6.1\%}} VideoMMMU, \textbf{\textcolor[HTML]{00FF00}{+4.8\%}} LongVideoBench), while facilitating \textit{cross-query knowledge reuse} and \textit{cross-model generalization}.