Knowing What is Missing: Efficient Conversational Memory via Explicit Evidence-Gap Tracking
Xingbo Du ⋅ Loka Li ⋅ Duzhen Zhang ⋅ Leonard Song ⋅ Le Song
Abstract
Long-term conversational memory poses a fundamental system challenge for LLM agents: reasoning over all past interactions improves coverage but incurs prohibitive token costs and noise, while retrieval-based methods are efficient yet often fail on multi-hop and temporal questions. Existing multi-step retrievers partially address this gap, but typically operate over an ever-growing textual history, causing context expansion and noise accumulation across iterations. We propose MemR$^3$, a backend-agnostic closed-loop controller that reformulates conversational memory retrieval as explicit stateful decision-making. At each iteration, MemR$^3$ maintains an evidence-gap state consisting of grounded information (evidence) and unresolved information requirements (gaps), and uses this state to route among three actions: retrieve, reflect, and answer. Newly retrieved snippets are incorporated into the state and then masked from subsequent prompts, so the model conditions on a compact summary of progress rather than the full retrieval history. This design turns multi-step memory retrieval into a bounded, inspectable control process whose query-time context scales with the state instead of raw accumulated text. Experiments on LongMemEval$_s$ show that MemR$^3$ surpasses implicit multi-step retrieval baselines and often approaches or exceeds full-context prompting while using only 1%-5% of the full-context tokens on long conversations. On LoCoMo, MemR$^3$ consistently improves over its underlying RAG- and Zep-based memory backbones. These results suggest that explicit evidence-gap tracking is an effective abstraction for token-efficient conversational memory.
Chat is not available.
Successful Page Load