Planning with the Views: Grounded Future-View Reasoning for Vision-Language Models
Abstract
Reliable embodied VLMs must ground sequential decisions in visual evidence and in the consequences of their own actions. We study this requirement through view planning, where a model moves a camera in a real 3D scene and localizes a target pose specified by its expected observation. We define a view world model as an action-conditioned representation of how observations change under viewpoint motion and introduce ViewSuite to evaluate both local transition reasoning and multi-turn use of that reasoning. Path-to-View and View-to-Path test forward and inverse camera transitions, while Interactive View Planning requires active observation acquisition with full 6-DoF control. Across 13 frontier VLMs, the best models exceed 70% on short-horizon transition tasks but reach at most 21.3% interactive success. Moreover, for the five analyzed frontier models, at least 90% of successful episodes occur only after an observation enters the target neighborhood. Current success is therefore usually coupled to search and visual matching rather than reliable anticipation of an unseen view. To provide grounded training signal, we retain every experienced action-observation transition, including those from failed trajectories, in a view graph. Distilling graph paths into the policy and alternating with self-exploration improves Qwen2.5-VL-7B from 2.5% to 47.8% interactive success. Behavioral analysis nevertheless shows that the trained model still largely approaches until the target becomes observable. ViewSuite thus separates improved evidence acquisition from faithful future-view imagination.