Planning with the Views
Abstract
Embodied spatial reasoning requires more than identifying spatial relations in observations. An agent must anticipate how its visual evidence will change as it moves and compose those changes into decisions in 3D space. We study this capability through a view world model, an action-conditioned representation of observation changes under viewpoint motion. We introduce ViewSuite, a benchmark built on 286 real ScanNet scenes with full 6-DoF camera control. Path-to-View tests forward reasoning from a camera-action sequence to its resulting observation, View-to-Path tests inverse reasoning between two observations, and Interactive View Planning requires an agent to choose a sequence of views and localize a visually specified target pose. Across 13 frontier VLMs, the strongest models exceed 70% on short-horizon forward or inverse transitions, but the best interactive planning success rate is only 21.3%. This exposes a composition gap between local camera-motion knowledge and goal-directed spatial reasoning over multiple views. We address sparse interactive rewards by accumulating all on-policy observation transitions, including those from failed episodes, into an action-labeled view graph. Graph paths provide grounded spatial trajectories that can be reformulated as dense supervision and distilled into the policy. Alternating this distillation with self-exploration raises Qwen2.5-VL-7B from 2.5% to 47.8% interactive success. The results identify future-view composition as a central bottleneck for embodied spatial reasoning, while reliable localization before the target neighborhood is observed remains open.