StepShield: When, Not Whether to Intervene on Rogue AgentsLLM-as-a-judge
Abstract
LLM judges are increasingly deployed as step-level safety evaluators that watch an agent act and decide whether to intervene. We show that the metrics used to validate such evaluators, accuracy, recall, and F1, are not construct-valid for the property that matters in deployment: whether the judge fires in time. We introduce StepShield, to our knowledge the first benchmark that scores when a monitor fires relative to an annotated divergence step, and define the Early Intervention Rate (EIR), the fraction of detected rogue trajectories flagged within a k-step window after the divergence point. EIR is informationally non-redundant with accuracy by construction. We release 9,429 incident-grounded code-agent trajectories with step-level divergence labels and evaluate four monitors on a 216-trajectory held-out split. EIR exposes the Forensics Trap: a pattern-based guardrail with 847 rules achieves 86% recall, yet its alerts carry no statistically detectable timing signal (EIR 0.23, not significantly better than a uniformly random step by a one-sided binomial test, p = 0.66), because 76% of its detections fire on benign prefix code before any violation occurs (mean intervention gap −5.0 steps). The 4× EIR gap between this guardrail and an LLM judge is invisible to F1. Judge validity is also a pipeline property: the same LLM judge that reaches EIR 0.89 with 5.6% FPR standalone drops to EIR 0.40 with 44.4% FPR inside a regex-first cascade, and its timeliness depends on being told what to look for. No evaluator we test approaches the ideal of 0% false-positive rate, 100% recall, and EIR 1.0 at once, and the hardest categories remain open.