The Selection-Target Problem in AI-Assisted Science: Reviewer Novelty and Later Scientific Uptake
Abstract
As AI systems generate more scientific ideas, selecting which ones merit attention becomes a central design problem. Reviewer-derived scores are attractive targets for this step. Yet reproducing a human score is useful only if it carries information about the downstream criterion the system is intended to serve. We instantiate a decision-conditioned target-validation test using Technical Novelty and Significance (TNS) and Empirical Novelty and Significance (ENS) from more than 3,800 ICLR 2022–2023 submissions. After conditioning on publication year and acceptance, TNS citation rank association falls from ρ=0.214 to 0.017 while ENS retains +0.192. This contrast persists across fixed citation horizons, identity controls and reviewer-level analyses. To examine what these reviewer scores may omit, a label-blind abstract-derived profile separately rates novelty claims, significance, evidence and clarity. Uptake-oriented dimensions retain association under the same controls and show stronger retrospective ranking association than reviewer novelty scores. Two raters agree on how the abstracts present significance, evidence and clarity. A post-cutoff cohort preserves the positive profile–citation direction under its unconditional design. Reviewer novelty is therefore not a unitary selection target: adjacent sub-scores can have sharply different relationships with later uptake. More broadly, AI-assisted science requires outcome-specific and decision-aware target validation before optimization. Because citations measure uptake rather than scientific value, the profile remains a diagnostic rather than a deployment metric.