Saliency-Aware Multi-Route Thinking: Grounding and Reasoning on Vision-Language Agents
Abstract
A vision-language (VL) agent wraps a frozen vision-language model (VLM) into a test-time system that may call external tools during execution before producing a response. The central problem is how the VLM and its tools interact across multiple steps. Two recipes are common, both rigid. Thinking for longer (borrowed from text-only agents) sharpens reasoning, but without in-loop visual refresh, it lets visual evidence decay into hallucination. Querying tools more often, as traditional agents do, treats outputs as ground truth, a rigid commitment that fails when calls are noisy. We propose Saliency-Aware Principle Selection (SAP), a training-free inference-time method that avoids both rigidities: it searches a broad space of high-level principles, short textual directives (e.g., ``re-examine the image whenever an intermediate conclusion is formed'') that each prescribe a different way for the VLM to weigh tool advice across steps, and selects the principle whose parallel reasoning routes best fit the actual visual evidence. Tool outputs are treated as advisory (consulted, never authoritative), and a population-based evolutionary loop drives the search. Under the same token budget as long chain-of-thought, SAP substantially reduces object hallucination while remaining competitive on reasoning-heavy benchmarks. Being training-free and composed of independent routes, SAP is plug-and-play on any VLM and parallel for modern agent deployment.