When Accumulated Memory Starts to Hurt: Forgetting and Cost in Non-Parametric Continual Adaptation
SAIKUMAR DANDLA ⋅ Harish Y V S
Abstract
An enterprise agent that adapts without retraining usually accumulates knowledge into an external memory and injects it at inference. This is attractive because it is cheap, inspectable and governable — and because, with no weights moving, it appears to avoid catastrophic forgetting entirely. We audit one such memory end to end and find that the forgetting reappears at the point of injection, that the memory does not survive a change of base model, and that the pipeline which built it silently corrupted its own contents. A multi-agent pipeline accumulated a Modern Standard Arabic trust-and-safety memory from its errors on a training split. Injected into the 3B generator it was built for, it raises AraTrust accuracy from 60.82% to 79.92% under a shared answer extractor (exact McNemar $p = 5.4 \times 10^{-7}$); a token-matched 100-shot control scores 45.81%, so this is not a context-window effect. The gain is not additive. Per item, injection moves 10 answers from correct to incorrect while recovering 7, and on 9 of those 10 the selected option changes — interference, not re-grading. The losses concentrate where the model was already competent: the two categories that lose without compensating gains are ones it answered at 100.0% and 89.7% unaided. The memory also fails across a base-model update, the ordinary non-stationarity of a deployed agent: applied unchanged to three larger generators it falls below what they reach with no memory on two of three. We propose but do not establish a mechanism — injection compresses accuracy toward the mean, so its value falls as competence rises (within-generator slope 0.452 rather than one; with two generators it clears no threshold under clustering-aware inference, $p = 0.21$–$0.22$). Two findings bear on whether such memories should be trusted. The accumulated memory loses to a fixed 18-token prompt on every axis we can measure: the prompt scores higher (90.45% vs 86.16%) at 63× fewer served tokens and 107 ms faster per call (paired, $t = 12.1$, $n = 171$), while the memory cost 41× more wall clock to build and never amortises. And it arrived partly corrupt: the pipeline embedded an answer sheet pairing seven training items with their answers, one entry fabricated — naming three options as the single correct answer to a medical question — and served to 153 of 171 test items. We release the memory as deployed and repaired, the router, the extractor, per-item predictions, and scripts reproducing every number.
Chat is not available.
Successful Page Load