AgenTracer-v2: Agentic Failure Tracer for LLM Agentic Systems
Abstract
LLM-powered agentic systems increasingly operate across complex, long-horizon task regimes, yet their growing architectural complexity has intensified system-level fragility and failure rates. The task of \emph{agentic failure attribution} seeks to identify the specific agent or execution step that causes failure within lengthy traces. However, most existing approaches remain constrained to passive, full-trajectory analysis, which limits their effectiveness in ultra-long and dynamically evolving trajectory attributions. To address this limitation, we introduce AgenTracer-v2, an autonomous, multi-hop failure attribution agent that reconceptualizes tracing as an active, tool-augmented diagnostic process. Concretely, we develop Tracer-Flywheel, a framework-agnostic data synthesis engine with deterministic system rollback, yielding an 8K dataset of precisely annotated failure trajectories. Leveraging this supervision, AgenTracer-v2 is trained to perform token-efficient global summarization, targeted inward inspection, and outward exploration via integrated diagnostic tools. Extensive experiments demonstrate that AgenTracer-v2 \textbf{(I)} consistently surpasses proprietary models such as \textsc{GPT-5.2} and \textsc{Gemini-3-Pro}, as well as prior specialized tracers, by up to 21.83% in attribution accuracy, \textbf{(II)} maintains robust performance on ultra-long trajectories exceeding 300K tokens and \textbf{(III)} delivering corrective feedback that measurably improves downstream agentic systems and multi-LLM reinforcement training.