LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems
Abstract
Agents in real-world settings operate over long and evolving horizons, where information is repeatedly updated and can interfere with each other across memories, requiring accurate recall and aggregated reasoning over multiple pieces of information. However, existing benchmarks focus on static, independent recall and fail to capture these dynamic interactions between evolving memories. In this paper, we study how current systems perform in realistic, continuously evolving, long-horizon settings across diverse domains and question types. To this end, we construct an analytical benchmark, LongMINT (Long-Horizon Memory under INTerference), which features (1) long, highly interconnected contexts with frequently updated information, (2) diverse coverage across multiple memory domains (Wikipedia, code, multi-turn dialogue, and state tracking), enabling evaluation of domain generalization, and (3) diverse question types, including (i) single-target recall tasks that test retrieval under interference over long contexts and (ii) multi-target aggregation tasks that require counting, ordering, or reasoning across multiple relevant pieces of information. We evaluate over six representative systems, including vanilla long-context LLMs, retrieval-augmented generation methods, and memory-augmented agent frameworks. We observe consistently low performance (avg. 27.7\% accuracy), especially on questions that require aggregated reasoning over multiple pieces of evidence. Fine-grained analysis shows that performance is primarily limited by retrieval and memory construction capabilities. Furthermore, current memory systems struggle to recall and reason over facts that are multiple steps back, and performance decreases when this lookback distance increases. These findings highlight the need for more robust memory management systems for dynamic, long-horizon environments across varying domains.