Silent Regressions: The Verification Gap in Coding Agents
Abstract
Coding agents in deployment are monitored by whatever is cheap to collect: did the run finish, did it return a patch. Running the project's tests would settle whether the patch is correct, but a deployed system can rarely do that for every proposed edit. We ask what the cheap signal misses and how little test execution it takes to recover, using 2,750 trajectories from four frontier models on 100 SWE-bench Verified tasks across 11 repositories, five runs per pair, joined to the harness evaluation logs so every failing patch can be labelled rather than merely counted. Completion barely separates agents that differ enormously on correctness: Claude and GPT-5 return a non-empty patch on 98.0% and 96.8% of runs, 1.2 percentage points apart, while resolving 67.4% and 45.2%, a gap of 22.2 pp [16.2, 28.4]. What completion hides is broken code rather than wasted effort. We call the failure a silent regression: a patch that breaks a test which was passing before the agent ran and that none of the non-test checks we evaluate separates from a fix. A randomly chosen Llama 4 run breaks a previously passing test with probability 40.0% [34.2, 46.2], and those patches apply cleanly and edit the same files as the reference fix. Two ways of avoiding test execution both fail. Prompting GPT-5 to write and run a failing test moves neither resolve nor regression rate, equivalent to baseline at a 10 pp margin; and a monitor over trajectory features predicts regressions on unseen repositories but falls to chance within a single model. On this evidence, deployment-time verification should rest on neither. What it should rest on is which tests, not how many. Running the benchmark-provided FAILTOPASS tests, supplied externally rather than written by the agent and 1.7% of the evaluated suite when tests can be selected individually, cuts the share of accepted patches carrying a silent regression from 18.8% to 4.8%; spending that identical budget on randomly chosen tests leaves 10.0%, twice as many. The saving comes from aim, and a verifier that cannot aim gets little from the same number of tests. Because most of the variation in whether a run resolves lies between tasks rather than between repeated runs of the same task (62–76%), every interval here resamples tasks rather than runs.