PIVOT: Pivotal-Decision Identification via Observable Structure
Abstract
LLM agents solve tasks through long sequences of tool-use decisions, but only a few of those decisions actually shape the outcome. Identifying the pivotal decisions is essential for auditing, debugging, and evaluating agent runs, yet the dominant approach (prompting a strong LLM as an impartial judge) is expensive, non-deterministic, and depends on the agent’s private reasoning. This is problematic because most production traces do not record it, and where it is recorded it can itself carry sensitive content. We present a deterministic framework that scores each decision from predominantly structural signals: what each action does (its role), how it follows from the previous step, problem signals, error cues, and position. Each signal is a pointwise-mutual-information (PMI) likelihood ratio. One of them, a transition signal, quantifies how unusual it is for the current action’s role to follow the previous one, using a matrix of role-to-role transition frequencies. The five signals are then combined into a single importance score per decision. Across two families of coding-agent traces, this empirical structural ranker matches or exceeds zero-shot LLM-as-judge baselines (Claude Haiku, Sonnet, and Opus): it substantially outperforms all judges when the agent’s recorded reasoning is limited and is on par with even the strongest judge when reasoning is fully available, at negligible cost, using no tokens and no GPU compute.