StepShield: When, Not Whether to Intervene on Rogue Agents
Abstract
Locating the step at which an agent's behavior went wrong is one concrete act of interpreting agent behavior, yet agent-safety benchmarks score only whether a monitor eventually flags a run, not when. Timing is the difference between intervention and autopsy. We introduce StepShield, to our knowledge the first benchmark that scores when a monitor fires relative to an annotated divergence step. We release 9,429 incident-grounded code-agent trajectories with step-level divergence labels and evaluate four monitors on a 216-trajectory held-out split. We define the Early Intervention Rate (EIR): the fraction of detected rogue trajectories where the alert fires within a k-step window after the divergence point, isolating timing quality from coverage. EIR exposes what we call the Forensics Trap: a pattern-based guardrail with 847 rules achieves 86% recall, yet its alerts carry no statistically detectable timing signal (EIR 0.23, not significantly better than a uniformly random step by a one-sided binomial test, p = 0.66), because 76% of its detections fire on benign prefix code before any violation occurs (mean intervention gap -5.0 steps). The 4× EIR gap between rule-based and semantic monitors is invisible to accuracy, recall, or F1. The failure is structural: regex guardrails match syntax, not intent, and therefore cannot locate the moment an agent turns rogue, which makes pattern-based monitors unsuited for real-time oversight in our setting. No method we evaluate approaches the ideal of 0% false-positive rate, 100% recall, and EIR 1.0 at once, and the hardest categories remain open.