When to Sleep: Consolidation Timing as a Stability-Plasticity Problem for Self-Improving Agents
Abstract
Self-improving agents adapt through an external experience memory and through their own weights. Position papers agree both are needed and call for a consolidation channel, but leave when to trigger it open. We make the trigger the experimental variable: holding the mechanism and the training-episode and reflection budgets fixed, we compare four fixed timing policies against a utility-triggered one (UTIL) on a five-environment tool-use stream with task boundaries but no test-time task identity (3 orders × 3 seeds), and repeat one comparison on LIBERO-Goal with a 450M vision–language–action policy. UTIL is the seven arms' only accuracy–cost Pareto point, dominating the rest: 49.7 ACC at 51% of memory-only's tokens per episode, backward transfer +1.1. It fires as often as the schedule it beats by 2.1 ACC points: when matters, not how often. Memory-only was cheap and stable, refuting our pre-registered hypothesis; weight updates raised constraint violations by up to 6.3 points, which ten replayed exemplars cut to 0.7; and the pretrained embodied policy barely forgets, so consolidation there buys storage rather than accuracy. Protocol, hypotheses, analysis code and runs are released.