EVA: Evidence-seeking Visual Agent for Hallucination-Resistant Multimodal Reasoning
Abstract
Hallucination remains a fundamental challenge in multimodal large language models (MLLMs), often stemming from unverified visual assumptions and reasoning processes that are weakly grounded in observable evidence. Existing approaches attempt to mitigate hallucination through self-reflection, multi-agent debate, or counterfactual verification. However, these methods remain largely opinion-driven: they lack explicit verification of the visual understanding process and may fail when multiple agents share the same incorrect perceptual assumptions, leading to consistent yet hallucinated conclusions. A promising direction is to augment MLLMs with external visual tools (e.g., OCR and grounding models) to acquire verifiable evidence. However, effective tool-augmented reasoning remains challenging due to two key failure modes: unreliable tool invocation, where models select inappropriate tools or produce incorrect parameters, and evidence--reasoning inconsistency, where models ignore or contradict the acquired evidence when forming final predictions. In this work, we propose EVA, an Evidence-seeking Visual Agent that reframes multimodal reasoning as an explicit process of evidence acquisition and verification. At the core of EVA is a unified evidence-grounded execution pipeline designed to address the above challenges through two tightly coupled mechanisms. First, we introduce \textit{Iterative Tool Replanning}, which dynamically refines tool selection and parameters based on intermediate observations, improving the reliability of tool usage. Second, we develop \textit{Evidence Consistency Verification}, an answer validation mechanism that improves alignment between the final prediction and the collected evidence, mitigating evidence--reasoning inconsistencies. In addition, to balance robustness and efficiency, EVA incorporates \textit{Selective Evidence Routing}, an adaptive fast--slow reasoning strategy that determines when external evidence is necessary. Visually straightforward queries are handled via direct reasoning, while complex or high-risk cases trigger tool-based evidence acquisition. By jointly addressing how to reliably acquire evidence and how to use it consistently, EVA transforms multimodal reasoning from passive answer validation into explicit evidence-backed decision making. Experimental results demonstrate that EVA achieves significantly improved reasoning reliability over direct VLM inference and multi-agent methods, while maintaining efficiency through selective tool invocation. These findings highlight structured and adaptive evidence utilization as a practical paradigm for hallucination-resistant multimodal reasoning.