Planning with the Views: Continual Self-Improvement through View-Graph Distillation
Abstract
Foundation-model agents are expected to improve through interaction, but sparse task rewards make most embodied experience appear unusable. We study a concrete continual self-improvement setting through view planning, where a VLM agent must choose camera actions and localize a visually specified target pose in a real 3D scene. ViewSuite separates local action-conditioned view-transition knowledge from the multi-turn use of that knowledge. Across 13 frontier VLMs, the strongest models exceed 70% on short-horizon transition tasks, yet the best Interactive View Planning success rate is only 21.3%, revealing a substantial gap between pretrained view knowledge and embodied action. To turn interaction into reusable experience, we retain every action-observation transition, including transitions from failed episodes, in an action-labeled view graph. The graph serves as persistent episodic memory and a structured view world model. Its connected paths are relabeled as goal-conditioned supervision and periodically consolidated into the VLM policy through distillation. Alternating on-policy exploration, graph expansion, and policy distillation creates a continual improvement loop: the current policy determines which experience is collected, memory organizes that experience, and the updated policy changes which regions become reachable next. Qwen2.5-VL-7B improves from 2.5% to 12.0%, 27.9%, and finally 47.8% across successive stages, while a graph collected through random actions reaches only 13.0%. The method also reaches 32.5% with Qwen3-VL-8B. These results demonstrate training-time self-improvement and knowledge consolidation in stationary reconstructed scenes, not test-time online learning, adaptation to nonstationary dynamics, or resistance to catastrophic forgetting.