Plan Shallow While You Learn: Reallocating Planning Compute During World-Model Adaptation
Abstract
A continual world model that adapts to a changed environment does not adapt instantly. Between the moment the physics shift and the moment the weights catch up, the deployed system plans inside a simulator that disagrees with reality, and its returns drop, often below what the stale frozen model would have delivered. We argue that the right response to this transient is not only to keep learning but to change how the planner spends its compute: move budget out of imagination depth, whose errors compound with every imagined step and therefore scale with how wrong the model currently is, and into optimization breadth. In a controlled deployment study over three drag shifts of a pursuit-evasion environment we find that (i) the profitable imagination horizon compresses with the severity of the shift; (ii) reallocating a fixed planning budget away from the healthy model's optimal depth wins on every shifted tier and loses on the unshifted control; and (iii) in live deployment streams the reallocated planner rides above the planner left on the default allocation, through the transient and after it. We propose that continual world-model systems treat planning-compute allocation as adaptive state, revised alongside the model's weights.