Beyond Conservation: Biological Language Models Poorly Capture Human-specific Variation
Abstract
State-of-the-art biological language models are trained on vast sequence corpora representing organisms across the tree of life, enabling models to learn representations of biological sequences that generalize across a wide range of evolutionarily disparate organisms. However, for deepening understanding of human biology and disease, a more human-centric evolutionary timescale is important. Human constraint captures loci that have only recently been subject to selection within the human lineage and may have functional consequence for human-specific traits and diseases. We show that across modalities (DNA, RNA, and proteins), state-of-the-art biological language models capture deep evolutionary conservation but are weak proxies of human constraint. We show that measuring models' abilities to distinguish functional variation by using Mendelian disease-causing variants largely measures how well models capture deep evolutionary conservation alone. On other datasets of common variants associated with both organismal and molecular phenotypes, conservation-based approaches perform poorly but are improved upon by explicit incorporation of metrics of human constraint. These results point to an unmet need to develop methods that model biological sequences not just across deep evolutionary divergence but also at a human-focused timescale.