MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing
Abstract
Long-term LLM agents must decide what to write into persistent memory before future queries are known, yet existing evaluations measure only final QA accuracy, entangling memory writing with retrieval and reasoning. We introduce MEMAUDIT, a package-oracle evaluation protocol that turns memory writing into a finite, auditable optimization problem with a certified ground-truth optimum under a fixed storage budget. Each evaluation package specifies an experience stream, candidate memory representations, and future-query requirements, enabling exact measurement of how well a memory writer preserves task-relevant information. We instantiate MEMAUDIT with a semantic coverage objective under storage and exclusivity constraints and compute exact optima via certified optimization. Across controlled and naturalistic settings, MEMAUDIT disentangles representation quality, validity preservation, and budget-aware selection effects that end-to-end QA cannot isolate. This enables, for the first time, principled and reproducible evaluation of memory-writing policies independent of downstream retrieval and reasoning.