EvoMM: Reinforced Self-Evolving Multimodal Agentic Memory
Abstract
Multimodal large language model (MLLM) agents increasingly operate over long histories of dialogues, images, and evolving facts, yet finite context windows and weak persistent memory cause forgetting, temporal inconsistency, and loss of precise past evidence. Existing multimodal memory agents either discard visual evidence by reducing images to captions or treat memory as a static store that is built once and never revised during reasoning; self-evolving agents further accumulate experience via prompting without tight coupling to retrieval and memory revision. Closing this loop with reinforcement learning additionally suffers from credit assignment difficulties over long, multimodal action trajectories. To address these challenges, we propose EvoMM, a reinforced self-evolving multimodal agentic memory framework that unifies memory construction, retrieval, and evolution within a single agentic loop. EvoMM maintains a mutable two-level memory of turn-level short memories and session-level long summaries, organized by a heterogeneous multimodal memory graph with temporal, semantic, image-co-reference, and keyword edges. An iterative PLAN→RETRIEVE→REFLECT→ANSWER policy then revises long memory on the fly and distills structured retrieval experience into a cross-query store for reuse. To optimize this policy, we further introduce PRR-GRPO (Planning-Retrieval-Reflection GRPO), which augments outcome rewards with gated, action-aware process rewards and performs dense credit assignment over a memory-action tree with dual-scale advantage normalization. Experiments on Mem-Gallery, MMLongBench, and LoCoMo show that EvoMM consistently outperforms strong text-only, multimodal, and agentic memory baselines, establishing a training paradigm for memory that reads, revises, and evolves with each query.