Genomic Pretraining Learns to Read Designed Enhancer Motifs, but Not Across Assay Domains
Yuan Liang ⋅ Siu Chung ⋅ Grant H Watson ⋅ Massimo Poesio
Abstract
Genomic foundation models are increasingly used to predict regulatory activity from DNA sequence, yet it remains unclear whether their gains reflect transferable regulatory knowledge or interpolation within a particular sequence library and assay. We treat enhancer sequence-to-function prediction as a verification case study for genomic pretraining, testing it at three levels: transfer under library-level distribution shift, the sequence features underlying this transfer using known designed motif locations, and generalization to an independent human primary-cell endpoint. Under leave-one-library-out evaluation, fine-tuned HyenaDNA consistently outperforms its architecture-matched random initialization and a 4-mer ridge baseline. This effect replicates with Nucleotide Transformer and in two additional independent MPRA studies, with a cross-study pattern consistent with a larger advantage in lower-data settings. The gain over a validation-selected, matched-capacity LegNet, however, is small. In-silico mutagenesis shows that pretraining improves localization of the designed motifs from an AUROC of $0.55$ to $0.66$. Across the tested sequence-model families, performance on a human erythroid primary endpoint remains low (Spearman $\rho\leq0.19$), whereas the measured discovery-assay readout predicts the same endpoint ($\rho=0.76$). Genomic pretraining therefore learns useful motif-sensitive features that transfer across designed sequence libraries, but these gains do not justify treating sequence-only models as stand-alone predictors of primary functional activity. A model can generalize across sequences while failing to generalize across the experimental measurement process.
Chat is not available.
Successful Page Load