Residue-Class Dependence of Twin Prime Gaps as Second-Order Singular Series Structure Discovered by a Machine Learning Pipeline
Wenhao Lu
Abstract
AI discovery pipelines produce claims faster than they can be checked, and the hardest failures are the ones that look verified: statistically enormous, internally replicated, and still wrong or mislabeled. We dissect a complete case from mathematics, chosen because its verifier is unusually strong, so every stage of the verification ladder can be climbed and audited. An ML pipeline run on all $3.4$ million twin primes below $10^9$ found a $\sim 20\sigma$ dependence of gap statistics on the residue class $p \bmod 210$, a signal that passed significance tests, range ablations to $10^{12}$, and symbolic-regression distillation, and which we initially drafted as a new standalone conjecture. It was not one. Deriving the class-conditional prediction of the Hardy-Littlewood pair-correlation singular series showed the "anomaly" is real but is second-order structure of a known conjecture, matching the data at $r^2 = 0.994$; a preregistered out-of-sample test on 135 unmeasured classes mod 2310 then confirmed it at $r^2 = 0.981$. The discovery survived, its interpretation did not, and the difference was invisible to every internal check the pipeline could run. We use the case to argue for a concrete verification ladder for AI-generated science (internal replication, ablation, distillation, derivation against theory, preregistered prediction), to locate which rungs are automatable, and to identify "novelty verification", deciding whether a true claim is actually new, as the step current AI scientist systems are least equipped to perform.
Chat is not available.
Successful Page Load