Long-Lived AI Agents Age Too: They Quietly Decay After Deployment
Abstract
The reliability of any long-running system degrades over time: databases accumulate stale indices, software accrues technical debts, and human memory fades with age. Memory-enabled agents are no exception: Even with frozen weights, their system state continues to change as they accumulate context across sessions, and their ability to store, retrieve, and apply knowledge deteriorates in ways that standard snapshot evaluation cannot capture. Recent benchmarks have begun to measure static degradation over long-horizon tasks, yet they do not diagnose what kind of degradation occurs, where in the agent memory architecture it originates, or how routine operational events reshape it. In this work, we introduce AgingBench, a longitudinal reliability benchmark suite organized around four aging mechanisms: compression aging, interference aging, revision aging, and maintenance aging. To localize aging, we perform component-level attribution via counterfactual analysis, identifying whether degradation originates in the write, retrieval, or utilization stage of the agent’s memory pipeline. We evaluated across 7 scenarios, 14 models, various memory policies and agent frameworks, ranging from a fully controlled runner to practical automatic agents such as Claude Code. Over ~400 runs across 8-200 sessions, we find aging is not one-dimensional: it can be invisible to behavioral tests while silently decaying; structurally sharp, with a single model shifting from perfect to zero accuracy on derived-state tracking; and stage-dependent, as strong models fail not at writing but at reusing their memory. Our findings demonstrate that as agents take on longer operational lifetimes, understanding the internal structure of how they age is as important as measuring how well they perform on day one.