When Conservation Prediction Does Not Transfer: Failure Modes in Allelic Variant Ranking
Abstract
Models trained to reproduce evolutionary conservation are often evaluated by their ability to rank causal variants, although direct annotation lookup, regional target prediction, and neural reference-versus-alternate scoring are different estimands. We reconstructed a 22M-parameter regional conservation model and audited two TraitGym scopes: a custom held-out-tile subset and the exposed public all cohort used for protocol parity. The model reproduced two positive-phyloP targets on 1,937 regional test windows with Pearson correlations of 0.803 and 0.770. On the custom subset, its chromosome-row-weighted AUPRC was unresolved against direct phyloP-241m for Complex variants (0.213 versus 0.219) and had a much lower Mendelian fixed-row point (0.388 versus 0.679). The 54 Mendelian positives represented 45 exact sites, 18 positive-bearing 1-kb loci, and 16 diseases; two diseases supplied 33 positives. A historical compact phastCons-versus-phyloP fixed-run difference of +0.081 became unresolved when 1-kb loci or diseases were resampled and changed to -0.002 after deleting the two dominant diseases. A one-base-versus-128-bp-mean comparator difference of +0.244 showed the same source-unit sensitivity. A discovered 3'-UTR advantage did not replicate in nonoverlapping follow-up; its point estimate changed sign, and direct phyloP coverage was nearly complete, leaving alignment-poor utility untested. The result is a failure analysis of evaluation transfer: strong regional prediction can coexist with unstable allelic conclusions when estimands, exposure, readout, and biological replication units are not separated.