Mechanistically Explaining the Pathogenicity of 2.1 Million Missense Protein Variants through Protein Language Model Representations
Abstract
A missense variant changes a single amino acid in a protein, and a central problem in genomics is predicting whether that change is harmful. Existing variant-effect predictors can rank substitutions by pathogenicity, but they often provide limited insight into what changed in the protein and why, which is critical for informing what subsequent experiments are necessary to achieve targeted therapeutic development. We address this gap by probing the learned, frozen evolutionary representations of a state-of-the-art protein language model, ESM-C 6B, to (1) predict the pathogenicity score for missense variants more accurately than state-of-the-art tools such as AlphaMissense [5] and (2) develop concrete insights into what properties of a variant are disrupted, including those that a genomic language model cannot provide. We find that covariance probing of frozen ESM-C representations ranks the pathogenicity of variants well, reaching 0.949 AUROC on a clinical missense benchmark, slightly above the same probe applied to the genomic model Evo 2 (0.945) and above AlphaMissense (0.943). To form a mechanistic understanding of variant disruption, we train more than 150 per-residue annotation probes across 22 million residues and 4,700 species that are able to recognize properties such as catalytic and binding sites, structure, topology, domains, motifs, and post-translational modifications. By measuring how these readouts shift under a substitution, we can suggest testable hypotheses about what was disrupted. Because the readouts are protein-native, they recover mechanisms a matched Evo 2 atlas misses. For instance, for the Parkin variant C441R, the binding-site readout from ESM-C collapses in step with the known loss of zinc coordination, while the genomic readout stays flat. Across 329,266 substitutions from 93 deep-mutational-scanning assays spanning 79 genes, MAPS pathogenicity scores correlate positively with experimentally measured functional damage (median assay-level Spearman ρ = 0.4864). We release these results as MAPS (Mechanistic Atlas of Protein Sequences), a queryable atlas that computes pathogenicity scores and mechanistic readouts for over two million missense variants across roughly 19,000 proteins, complementing the nucleotide-level view that genomic models provide. Code: https://anonymous.4open.science/r/maps-anonymous-0BC2. Interactive atlas: https://maps-atlas.pages.dev