Diagnosing Apparent Temporal Forgetting in Long-Term Memory Agents
P Sivadhanushya
Abstract
Long-term memory agents are often said to ``forget'' as relevant information recedes further into the past. Yet in the agents we study, benchmark calendar age is evaluator-side metadata and is never exposed to the model, so an age--performance association alone cannot identify temporal forgetting. We therefore localize long-horizon memory failures into three observable stages: memory availability, evidence recovery, and answer synthesis. Across Gemini 2.5 Flash and Qwen3-235B in a shared LongMemEval-derived memory-agent architecture, expanding the visible summary budget eliminates availability failure entirely, but converts only 17.9\% and 21.4\% of previously inaccessible cases, respectively, into success; most failures instead migrate downstream to incomplete evidence recovery. Availability is strongly associated with correctness in both models (OR 6.35 and 6.63). Under full visibility, a post-hoc continuous structural-position analysis over all 500 cases provides no robust cross-model support for a simple recency account. A complementary controlled intervention that relocates a gold-evidence session from the front to the back while holding calendar dates fixed likewise produces no statistically reliable recent-position advantage in complete evidence recovery (Gemini $-0.008$, 95\% CI $[-0.039,0.022]$; Qwen $+0.009$, CI $[-0.029,0.050]$). Thus, in this two-model case study, the remaining observable failure signature lies before complete annotated-gold recovery even when evidence is available. More broadly, temporal degradation curves should not by themselves be interpreted as evidence of forgetting: memory evaluations should distinguish evaluator-side time from agent-visible availability, retrieval, and downstream answer generation.
Chat is not available.
Successful Page Load