Which ClinVar Labels Can a Genomic Foundation Model Actually Learn From? Review-Status-Stratified Evaluation of AlphaGenome on Neurodevelopmental-Disorder Variants
Abstract
Genomic foundation models such as AlphaGenome are increasingly evaluated by their ability to separate pathogenic from benign ClinVar variants, and high aggregate separation is often read as evidence that the model has learned biologically meaningful pathogenicity signal. This reading treats every ClinVar label as equally trustworthy ground truth, but ClinVar attaches an explicit review-status field to every record, and single-submitter records are known to carry more label noise than multi-submitter, expert-panel records. We audit this data-readiness question directly on a balanced 15-variant pilot of neurodevelopmental-disorder (NDD) ClinVar variants scored with AlphaGenome's 99th-percentile absolute quantile statistic: mean scores track clinical significance in the expected direction (0.9951 pathogenic, 0.9906 VUS, 0.9822 benign), but 13 of 15 records carry single-submitter review status and the pilot contains zero expert-panel or practice-guideline records, leaving only two multi-submitter records with which to even attempt a review-status comparison. This is itself a data-readiness finding: at the scale most pilots are run, NDD ClinVar material is too thin in its high-confidence review-status strata to support the stratified evaluation the field needs. We contribute (i) a validated, metadata-complete NDD ClinVar dataset inventory — coordinate-checked and annotated with review status, submitter count, and conflict flags — as a reusable data-readiness artifact, and (ii) a concrete, pre-registered protocol for scaling this inventory to a gene-capped, review-status-diversified panel (target N=180) and re-running the pathogenic-versus-benign comparison stratified by evidence tier. We report this as work in progress: the stratified re-scoring has not yet been run, and we argue that reporting this absence honestly, together with the artifact and protocol needed to close it, is more useful to the community than a pooled separation number that cannot say which evidence tier is driving it.