Two Faces of Missing Mass: Hallucination and Discovery in Language Models
Abstract
Large language models (LLMs) hold much promise for scientific discovery, yet basic questions remain about when they can reliably make discoveries. We study this question through the lens of \emph{missing mass}: the probability assigned to outcomes unseen in a finite sample. We extend recent missing-mass theories of hallucination~\citep{kalai2024calibrated} to include discovery. In particular, we show that monofacts--statements appearing exactly once in the training data--can bound discovery capacity. We test this prediction in materials and regulatory genomics by varying the training set's monofact prevalence. Across two open-weight model families and multiple training and generation settings, higher monofact ratios consistently yield higher discovery rates. Our results suggest that the long-tail of scientific data is not only a source of hallucination risk but also a statistical resource for discovery.