Continual World Modeling via Sleep
Abstract
What should a long-lived agent keep learning? We argue for its world model rather than its policy: action-conditioned prediction is supervised by every experience the agent has, including failures, and a fixed planner turns the predictions into actions. We propose Sleep-CWM, which alternates wake, when the agent acts while a LoRA adapter absorbs new experience, with sleep, when the base model consolidates new and replayed experience. A bounded buffer evicts the rare events that link a cue to its delayed consequence, so sleep keeps event and precondition windows outside it and replays whatever is being forgotten. A generative variant lets the pre-sleep model regenerate replay targets from stored cues. We evaluate retention, positive transfer, and compositional generalization in predictions and planner behavior. With a 0.7B video diffusion world model in Minecraft, Sleep-CWM retains hidden-consequence tasks on which a continually trained policy head fails and recovers what uniform replay loses under small routine buffers. Unlike a baseline that keeps the predictive score but not the rollouts, it gives the planner an early warning of hidden hazards.