FocusNav: Learning Task-directed Perception via Action-Aware Future Reconstruction in Vision-and-Language Navigation
Abstract
Vision-and-Language Navigation(VLN) requires agents to sustain task-directed spatial perception to focus on navigation-critical regions while translating linguistic intent into low-level actions. Prevalent hierarchical planning methods rely on a high-level policy to select waypoints or subgoals for low-level control, but they often fail to remain focused, causing inaccurate or redundant subgoals that directly mislead execution. Moreover, their separation between visual perception and action execution neglects fine-grained spatial reasoning in low-level movements and limits real-world deployability. To address these two issues, we propose FocusNav, a grounding-driven end-to-end monocular framework that learns task-directed spatial perception through action-aware future reconstruction. Concretely, FocusNav first boosts the agent’s visual attention on task-relevant regions with a Grounding-driven Visual Reconstruction module that reconstructs instruction-relevant targets and traversable areas via a semantic-guided latent diffusion process. Subsequently, to bridge the perception–execution gap, we incorporate physical action priors via Action-aware Visual Conditioning and enforce Intent-Effect Alignment between anticipated visual changes and observed action effects. These effectively enhance the agent’s task-directed spatial perception to direct attention to navigation-critical regions, and tighten the perception–execution loop. Experiments on R2R-CE and RxR-CE benchmarks show that FocusNav achieves substantially improved state-of-the-art performance with a simple monocular low-level policy, while providing faster inference than recent methods. Furthermore, real-world robot experiments validate our FocusNav’ s effectiveness, ease of deployment, and lightweight design.