When Good Rankings Make Bad Instruments: Cutoff Transfer Failure in AI Variant-Effect Scores
Abstract
AI systems increasingly enter scientific and clinical arguments as evidence, yet the benchmarks certifying them are rarely examined for what they measure. We take one such instrument as our object of study. Models that score DNA variants for harmfulness are judged by how well they rank harmful variants above harmless ones. A laboratory does not use a ranking; it picks a score threshold, or cutoff. We experiment with 40,976 clinically labeled human variants scored by 5 models, where these two properties come apart. Splitting the genome by evolutionary conservation leaves ranking quality unchanged, but a cutoff tuned to miss 5% of harmful variants where curated labels are dense misses 72.6% where they are sparse, and 92.2% for one long-established tool. Class proportions do not explain this, and two unrelated ways of splitting the same data leave the cutoffs intact, so the failure is tied to conservation, the signal these models are built on. The standard metric reads only ordering and so cannot see it: ranking improves in the stratum where the cutoff collapses. A benchmark number certifies a specific use, and reading it as general competence is a measurement error rather than a modeling one.