What DNA Foundation Models Learn Beyond Sequence Composition
Vincenzo Y. Civale ⋅ Andrew Bagdanov ⋅ Alberto Magi
Abstract
A central question in genomic representation learning is whether DNA foundation models~(FMs) encode information beyond simple sequence composition, or whether their apparent representational power largely reflects signals already captured by classical compositional statistics. We address this by establishing $k$-mer frequency vectors as a rigorous biologically grounded baseline and introducing a geometric decomposition that partitions FM embeddings into a component linearly predictable from $k$-mer frequencies and an orthogonal residual representing genuinely non-compositional information. By evaluating each component separately as a downstream feature space, we directly measure what FMs encode beyond composition and whether that content is beneficial, neutral, or detrimental. Across 57 genomic classification datasets and 2 gene expression regression tasks, we compare five FMs (NTv3-650M, HyenaDNA, DNABERT-2, Caduceus-Ph, Evo2-1B) against $k$-mer features ($k \in \{4,5,6\}$) under a controlled frozen-embedding protocol with a fixed Random Forest classifier. NTv3 is the only model to consistently outperform $k$-mers (63.2\% of datasets, mean MCC $= 0.562$), with gains concentrated in splice-site and promoter tasks; HyenaDNA, DNABERT-2, and Evo2 are outperformed on 86\%, 95\%, and 93\% of datasets. Wilcoxon signed-rank tests with FDR correction confirm these patterns are systematic and robust to dataset heterogeneity. Our decomposition reveals that FM advantage originates exclusively from the non-compositional residual, and only in tasks with inherently positional regulatory signals. For HyenaDNA and DNABERT-2, the residual actively degrades classification. The fraction of embedding variance explained by $k$-mers ($R^2$) is uncorrelated with downstream quality, demonstrating that the quantity of non-compositional information is not a proxy for its utility. These results show that, in the frozen-embedding regime, current DNA FM representations largely recapitulate sequence composition, and provide a principled diagnostic for predicting which task conditions are likely to benefit from non-compositional structure.
Chat is not available.
Successful Page Load