MEME: Multi-Entity & Evolving Memory Evaluation
Seokwon Jung ⋅ Alexander Rubinstein ⋅ Arnas Uselis ⋅ Sangdoo Yun ⋅ Seong Joon Oh
Abstract
LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. While prior benchmarks evaluate only single-entity updates, MEME defines six tasks spanning the full space defined by the multi-entity and evolving axes, including three not scored by prior work: Cascade and Absence (dependency reasoning) and Deletion (post-removal state). Evaluating six memory systems spanning three memory paradigms on 100 controlled episodes, we find that all systems collapse on dependency reasoning under the default configuration (Cascade: 3\%, Absence: 1\% in average accuracy) despite adequate static retrieval performance. Prompt optimization, deeper retrieval, reduced filler noise, and most stronger LLMs fail to close this gap. Only a file-based agent paired with Claude Opus 4.7 partially closes the gap, but at $\sim$70$\times$ the baseline cost, indicating closure currently depends on configurations that are not practical at scale. Code is available at https://anonymous.4open.science/r/MEME-0612 and the dataset at https://huggingface.co/datasets/meme-benchmark/MEME.
Chat is not available.
Successful Page Load