DoG: Sniffing Out Overconfidence in LLM Agents via Post-hoc Trajectory Restructuring
Abstract
The overconfidence of Large Language Model (LLM) agents poses a critical challenge to their reliable deployment for complex, high-stakes tasks. Current confidence estimation methods predominantly treat agent execution as a flat, one-dimensional, time-ordered sequence. This oversimplification fundamentally fails to capture the complex logical dependencies, branching paths, error propagation, and tool interactions inherent in agentic reasoning. To address this, we introduce DAG of Grounds (DoG), a novel method for agent confidence estimation via post-hoc trajectory restructuring. By transforming linear trajectories into Directed Acyclic Graphs (DAGs), DoG explicitly maps the grounding of an agent's final answer to its intermediate reasoning steps and tool outputs. We evaluate DoG on several benchmarks, GAIA, GPQA, and HLE, which involve long-horizon reasoning and iterative tool use, and show that DoG improves calibration performance across standard metrics such as ECE, Brier score, and AUROC. DoG outperforms existing calibration methods designed for general LLM outputs, performs comparably to or better than methods specifically designed for agent trajectories, and remains robust across different LLM backbones and tool use settings. Supported by an interactive diagnostic tool, our framework provides unprecedented interpretability for diagnosing opaque failure modes. Together, these contributions establish a structured, graph-based paradigm for confidence estimation that successfully "sniffs out" overconfidence in autonomous systems.