InfoNav: A Unified Value Framework Integrating Semantic Relevance and Information Gain for Zero-Shot Object Goal Navigation
Abstract
Zero-shot object goal navigation (ObjectNav) requires an embodied agent to locate a target object in an unseen environment without any task-specific training. Recent methods leverage large vision-language models (VLMs) to inject semantic priors, but still suffer from two structural issues: (i) semantic cues are treated uniformly across space, ignoring how the spatial extent and hierarchy of contextual entities modulate target likelihood; and (ii) semantic exploitation and geometric exploration are decoupled, producing brittle behavior whenever semantic cues are weak, conflicting, or absent. We address these issues by reformulating navigation as a unified value estimation problem over frontier candidates. Specifically, we propose InfoNav, which jointly models (1) a hierarchical semantic value map that instantiates a spatially-weighted total-probability decomposition of target likelihood through a multi-level spatial influence field; (2) a VLM-based estimator that approximates the intractable long-horizon information gain by grounding it in visual-semantic connectivity; and (3) a memory-augmented verification mechanism that replaces fixed confidence thresholds with a temporally adaptive schedule with deferred re-examination of uncertain detections. InfoNav establishes the highest zero-shot success rate on all three major ObjectNav benchmarks: 62.0% on HM3Dv1 (+0.6 over BeliefMapNav), 81.5% on HM3Dv2 (the only method above 80%), and 41.8% on MP3D (+4.5 over BeliefMapNav), while requiring more than 96% fewer multimodal-model calls than representative LLM-driven baselines. Extensive ablations confirm that the semantic and information-theoretic value fields are complementary, and that the unified formulation is what enables the agent to remain decisive under strong cues and exploratory under weak ones.