Turning AI-Generated Hypotheses into Testable Distributions: A Framework for Bayesian Verification
Abstract
AI scientists now generate hypotheses faster than they can be verified; the bottleneck of AI-driven discovery is shifting from generation to verification. Comparing a language-expressed hypothesis with data requires a statistical object. When an appropriate model family is known, expert-built mechanistic or probabilistic models fill this role, yet constructing one model per candidate does not scale with automated hypothesis generation. In place of bespoke per-hypothesis modeling, we train a hidden-state-conditioned flow on curves elicited from a frozen LLM teacher, distilling each hypothesis into a reusable distribution that supplies fresh samples and a numerical density. These distributions plug directly into Bayesian evidence, posterior hypothesis probabilities, and posterior-weighted prediction. A fatigue-crack-growth positive control verifies the downstream machinery -- hypothesis selection and posterior-weighted terminal-extent prediction -- on a language-defined hypothesis bank. In the adsorption study, sequential observations update five hypotheses that GPT-5.6 Sol proposed before any measurement was seen, yielding hypothesis evidence and posterior predictive distributions. Relative to a matched Bayesian monotone-spline baseline, the posterior-weighted mixture predicts each isotherm's unobserved remainder with 20--27% lower macro CRPS and RMSE after eight and thirteen of its 21 points, while selecting hypotheses consistent with the observed shapes, confirming that the end-to-end approach is operational. Improving local density transfer to unseen hypothesis texts, establishing overlap-based abstention diagnostics, and broadening the evaluation remain the next priorities; we additionally provide a reporting checklist for verifying AI-generated hypotheses.