Your Research Loop Is Measuring Noise: Accept Rules for Autonomous Experimentation
Brian Jalaian ⋅ Venkat R Dasari
Abstract
Autonomous research loops now run unattended: an agent proposes a change, trains, compares against the incumbent, and keeps the change if the number improved. The deployed autonomous-research systems we surveyed decide on a \emph{single} run; accept rules with error guarantees exist in the bandit and racing literature but are not wired into these loops. Whether that is sound depends on one ratio, $\rho$, between the effect the loop is hunting and the standard deviation of re-executing the same code. We measure $\rho$ for a production LLM distillation-and-quantization pipeline: every one of 16 conditions re-executed to $n=30$ under a pinned evaluation environment, 480 runs in all, with 1280 hyperparameter trials logged across the 80-run sub-block that also records validation loss. Replicate standard deviation spans more than two orders of magnitude across conditions; the effect this pipeline's authors were chasing sits at $\rho \approx 0.22$; and the cheap objective the search minimises points the wrong way: lower validation loss predicts \emph{lower} benchmark accuracy, where a working proxy would give the opposite sign ($r=+0.33$, $p=0.008$). Replaying accept rules over those real runs, best-of-one keeps 49.9\% of candidates that are in truth exactly neutral, peeking after every run inflates the false-accept rate to 11.3\% against a nominal 5\%, and a Robbins normal-mixture e-process stays at 5.3\% while remaining valid at every stopping time, the guarantee an autonomous loop needs, since it decides for itself when to stop. The rule's detection power tracks the measured noise floor: the same fixed accuracy gain is a large-SD effect on a tight condition, confirmed in a handful of runs, and a sub-SD effect on a noisy one, where no bounded-error rule can confirm it. Across all 16 conditions at $n=30$, estimating $\sigma$ from a five-run calibration phase inflates the error to 12.1\%, and the shortcut of reusing the incumbent's own runs is no better (11.9\%); only an oracle $\sigma$ (4.2\%) or a dedicated calibration phase of 20 runs holds nominal, a budget that survives leave-one-condition-out validation over all 16 conditions. The noise floor must be measured deeply, not estimated from a handful of runs. Finally we exhibit the failure in a published manuscript whose complete revision history was available to us, so we can date every column: its headline claim compares the maximum of five search replicates against one baseline run, the selection bias averages 1.67 points, larger than the claimed effect, and the column that carries it entered the paper only at revision, in answer to a reviewer asking for better statistics.
Chat is not available.
Successful Page Load