Automata from Agent Traces: Failure and Next-Step Prediction
Abstract
LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, missing the cross-run topology that links next-step and failure prediction; we extract provably minimal finite-state machines (FSMs) via prefix-tree construction and structural state merging, providing a structural substrate for the unpredictable nature of LLM agent behavior. Across twelve public datasets, the FSMs are compact (7–43 states), achieve high replay fitness on held-out data with zero structural variance across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context beats Agent Workflow Memory on all eight ground-truth-matched datasets. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor flags failing SWE-agent runs at rank-AUROC above the trivial flag-everything baseline, enabling early-stopping at 32% trace completion. A single FSM replays four LLMs at perfect fitness, evidence that behavioral topology in LLM agents is shaped more by the deployment harness than by the LLM and giving a model-agnostic structural primitive for safety auditing and runtime monitoring at deployment.