The Placebo Improvement: Stress-Testing False Claims by Autonomous Research Agents
Marek Suppa ⋅ Jaroslav Kopcan
Abstract
Benchmarks for autonomous research agents count the improvements they find, but rarely measure false claims: reports of gains when the code had no effect. We introduce a stress test for this failure mode. Each agent evaluates a plausible code change that we verify produces exactly the same outputs as the baseline on 100 reference seeds. A controlled executor hides the seeds and selects lower-scoring baseline runs and higher-scoring modified runs, creating an apparent gain without changing the underlying computation. We compare an honesty-oriented protocol that permits abstention with a deployment-like protocol that makes improvement the objective and requires a yes/no decision. The pipeline, interventions, reference set, and seed-selection rule remain fixed across protocols, although agents may request different runs and therefore receive different result rows. At a selected apparent gain of $+2.5$ percentage points, all nine configurations tested under both protocols make more false improvement claims under the deployment-like protocol. For Claude Opus 4.5, the null-claim rate rises from 16\% to 96\%. Five configurations with a 0\% rate under the permissive protocol reach between 12\% and 59\% under the deployment-like protocol. A low null-claim rate can also reflect abstention. Among configurations between 0\% and 3\%, the share that commits to yes or no ranges from 3\% to 84\%. GPT-5.6-sol committed in only 6\% of null episodes, claimed one of three scorable positive-control episodes, and produced no usable verdict in the other three. These are conditional stress-test rates, not estimates of a deployed program's base rate. Benchmarks should therefore report null-claim, decision, and positive-control rates alongside discovery counts.
Chat is not available.
Successful Page Load