Planning with View World Models
Abstract
World models should improve as agents gather new experience, yet sparse task rewards leave most interaction apparently unusable. We study continual model expansion through view planning, where an embodied agent must acquire observations and localize a visually specified target pose in a real 3D scene. Every camera action produces a valid transition between observations, even when the trajectory fails its original goal. We retain these transitions in an action-labeled view graph, a discrete view world model that accumulates across episodes. Connected graph paths provide grounded future-view rollouts and can be relabeled as new goal-conditioned planning examples. We alternate this graph distillation with further on-policy self-exploration, allowing the model and policy to co-evolve during training: the current policy determines which view space is explored, accumulated experience supplies denser supervision, and the updated policy reaches new observations. On ViewSuite, this procedure improves Qwen2.5-VL-7B from 2.5% to 47.8% interactive view-planning success. Performance rises across graph-distillation rounds, while a graph collected by random actions reaches only 13.0%. The results show that both structured retention and policy-directed acquisition matter for learning from self-generated experience. Our stationary reconstructed scenes study continual expansion during training, not changing dynamics, catastrophic forgetting, or inference-time model updates.