Validation Protects Truth, Not Coverage: Tail Starvation in Scientific Proposal Priors
Abstract
Model-collapse studies examine tail erosion under recursive learning. AI-assisted science can create a related loop: models propose hypotheses, pipelines validate them, and selected findings enter future corpora. We show that validation can preserve the truth of what gets accepted while narrowing what gets proposed: costly or off-paradigm findings may not re-enter the corpus, and some families of hypotheses may stop being generated. We model this loop with four steps: propose, validate, re-enter, and learn. In a simplified two-paradigm setting, we derive a boundary in the model's cost, value, and popularity effects above which a rare-but-valuable paradigm is driven from scarcity to extinction, even as average validated quality stays flat. In a symbolic equation-discovery system, the learned proposal strategy sharply down-weights the operators that rare functional families require, even as the quality of re-entered fits remains near-perfect. The same pattern holds on a fixed set of Feynman equations, and persists even when we remove the biases our model targets, pointing to an additional structural cause. We therefore separate the theoretical mechanism from this broader pattern of narrowing. Because average quality can stay high while coverage narrows, monitoring validated quality is not enough: proposal diversity and coverage must be tracked directly.