The Optimal Memory of Failure Is Lossy: Wrongful Convictions in Shared Failure Banks for Autonomous Research
Abstract
Autonomous research systems increasingly share negative knowledge: failed attempts are written to a common bank so that no agent repeats them. We show this design has a structural flaw. Failure verdicts are statistical claims produced by noisy, small-n experiments, and their two error types are asymmetric by construction: a false confirmation is built upon and self-exposes, while a false conviction is locked precisely so that nobody retries it. In a model of cumulative research where confirmed ideas beget new ideas, we prove that with exogenous supply and binary records full deference is optimal, but with success-gated supply discovery is inverted-U in deference: full sharing forfeits up to 81% of discoveries at long horizons, governed by the extinction probability of the idea-branching process. Measured wrongful-conviction rates track the closed form across three landscapes: 10–20% in two live training loops, a 50% point estimate for verification-stage near misses (the most evidence-laden negative verdicts are the least reliable), and 2.8% aggregate on a marginal-sparse public benchmark whose near-marginal configs are still convicted 44% of the time (theory: 45%)—the rate is set by the mass of marginal effects, not by any aggregate noise ratio. A behavioral study of twelve deployed LLMs across nine model families shows the lock is behavioral and replicates across families: free-choice revisits of DEAD-marked ideas are 0/94 (one-sided 95% upper bound 3.1%), and the label alone suppresses selection of an a-priori-best idea in 9/9 tested families (two-sided exact sign test p=0.0039). Belief responses, by contrast, are heterogeneous—provenance-unlockable, hard-locked, and calibrated-but-choice-locked (one family prices every condition at the Bayesian posterior yet still refuses the idea)—motivating record-level calibration as the common intervention point rather than belief repair. In end-to-end closed loops with real training runs, evidence-graded memory delivers 24–35% more truth-validated discoveries than binary locks in the noisy, discovery-rich regime, and cedes its edge in a low-noise cell exactly where the phase structure predicts. The fix is not less sharing but calibrated conviction: evidence-threshold locking dominates binary locks and probability matching, and is robust to 2× misspecification of every parameter it needs.