Assessing Novelty Without Ground Truth: Proxy Measures for Evaluating AI Scientists in Global Health
Abstract
AI scientists can generate candidate hypotheses at a rate far higher than any research team can evaluate them, with particular consequences for global health research. Often the teams closest to the disease burden have the fewest people, the shortest funding cycles, and the least room to invest in an idea that may turn out to be unviable. When evaluating AI-generated hypotheses, the core difficulty arises when considering scientific novelty, as the originality of a generated hypothesis has no ground truth to verify against. A hypothesis significantly distinct from existing literature could reflect a genuinely original connection, or it could simply rest on unsound claims. We argue that verification and proxy measures for novelty must be assessed and reported separately, as a composite score that conflates them prevents a domain expert from distinguishing one case from the other. We propose a framework that enforces this separation, giving resource-constrained teams a principled basis for deciding which hypotheses are worth pursuing.