A Frozen Model That Keeps Improving: Label Vintage and the Validity of ClinVar-Based Evaluation
Abstract
Genomic AI models report their headline accuracy against ClinVar, the clinical variant archive. But ClinVar's labels are not observations of nature. They are expert judgements produced under the ACMG/AMP framework, which explicitly admits computational predictor output as evidence through the PP3 and BP4 criteria. The answer key is therefore partly written by the class of system being graded against it. We ask what that does to the benchmark, and propose a diagnostic that requires no new model and no new labels: the frozen-clock test. A model whose weights never change has a fixed relationship to biology, so if its measured accuracy varies with the vintage of the labels it is scored against, the labels moved, not the model. Applying this to 188,523 ClinVar missense variants scored by two frozen predictors, we find that REVEL—the tool ClinGen calibrated for PP3—gains +0.0039 AUROC per year of label vintage while AlphaMissense gains +0.0017, on the same variants, holding gene and review status fixed. The paired gap moves from +0.004 to −0.015 and lies far outside a permutation null. The consequence is that the reported ranking of the two models is a property of label vintage: REVEL leads on ClinVar overall and on post-2022 labels, while AlphaMissense leads on labels first classified before 2018. We also report what did not survive scrutiny: a reclassification effect that looked strong reversed sign under matching; an evidence-provenance contrast that was dramatic in AUROC shrank substantially under an imbalance-free metric; and the apparent growth of computational evidence in ClinVar rationales is largely one submitter. We argue that benchmark reports for scientific AI should state the vintage and evidentiary provenance of their labels, and that the frozen-clock test is portable to any benchmark whose ground truth is produced by a human process that has itself adopted AI.