MemContract: Contract-Sensitive Evaluation for Mutable Agent Memory
Abstract
Many agent-memory evaluations collapse state revision, deprecation, erasure, and provenance into retrieve-heavy scores. We introduce MemContract, a diagnostic benchmark of 1 , 080 multi-session tasks over store, retrieve, revise, deprecate, erase, and prove-origin, with formal pre/post-conditions, mandatory distractor sessions plus alias rotation on every family, blinded/stale/delayed controls, a fixed 25-probe residue audit, and benchmark-validity checks via a 180-task human contract-fit audit and a 120-task human-assisted upper bound. Under matched prompts, compute, storage, and latency budgets on GPT-4o across seven architectures, Graph Memory reaches 67.6% macro average (95% paired-bootstrap CI [65.8, 69.4]) versus 62.4% for the strongest hybrid baseline and 40.6% for RAG, with most separation concentrated in revise, deprecate, erase, and prove-origin while store/retrieve remain comparatively compressed. Claude 3.5 Sonnet and Gemini 1.5 Pro 002 preserve the same top-three ordering on a matched 360-task rerun, while the manual audit finds 96.7% agreement that the harness label matches the intended contract and the human-assisted upper bound reaches 93.8% macro, placing the best-system score in a solvable regime rather than a floor of label noise. Reviewer-critical calibration checks tell a narrower but durable story: removing prove-origin leaves the Graph - Hybrid gap unchanged at 5.2 pp, best-effort prompt/schema tuning narrows it to 2.6 pp, and fully external LongMemEval wu2024longmemeval and LoCoMo maharana2024locomo narrow it to 1.8 and 2.8 pp while preserving the top-three ordering Graph > Hybrid > Editable KV. The stable claim is therefore not universal backend superiority: contract-sensitive evaluation reveals separation that broader or retrieval-heavier workloads attenuate but do not erase. The 25-probe audit bounds catchable residue leakage at 3.9% for Graph Memory and 27.4% for RAG, but does not certify deletion.