Evidence Guided Adaptive Search for Helping AI Agents Find Missing Facts Before Answering
Akhil Kasturi ⋅ Min Cheng ⋅ Huilin Lu ⋅ Bingrou Zhou ⋅ Qiuying Lin ⋅ Ge Shi
Abstract
Large language model (LLM) agents have progressed rapidly in reasoning, tool use, and long-horizon planning, moving from simple question answering to systems that search, retrieve, verify, and act across complex environments. Recent work pushes this further by letting agents critique their own mistakes, revise their plans, and search over alternative reasoning paths before committing to an answer. Most of these systems, however, still treat a task as a single trajectory to be extended, scored, or corrected, and they struggle precisely where reasoning harder along one path cannot help. Many real-world tool-use tasks leave the agent without enough information to tell which of several plausible answers is correct. This implies that the agent begins a multi-hop question or a stateful workflow without knowing which facts will decide the answer, so several continuations look equally plausible until an observation rules the wrong ones out. The difficulty here lies in the missing information rather than in the reasoning, and existing methods carry no explicit representation of that gap. They either commit early to an unverified guess or spend further steps re-deriving the same unresolved reasoning in a new form, without going out to find the fact that is missing. In this work we treat a task as a set of unknowns to be resolved rather than a single path to be solved outright. We propose Evidence-Guided Adaptive Search (EGAS), an adaptive agentic search algorithm that decomposes a query into a structured set of unknowns, dispatches independent agents to resolve them in parallel, and records in a shared evidence ledger only what the observed tool outputs support. A lightweight controller reads that ledger, never an agent's own account of its reasoning, to decide whether to deepen an unknown, try a different angle, resolve a conflict between two recorded facts, or stop once the task is determined. On BFCL V4 Web Search and $\tau^2$-bench Retail, EGAS reaches the highest accuracy of the seven methods compared, improving over the strongest baseline in every condition and by roughly $8%$ on average across both benchmarks. This advantage holds across both backbone LLMs we tested, Claude Sonnet 4.6 and Claude Haiku 4.5, where EGAS is the most accurate method under both.
Chat is not available.
Successful Page Load