Interpreting Latent Protein Language Model Features with Geometric Annotations
Siddharth Setlur ⋅ Djordje Mihajlovic ⋅ Darrick Lee
Abstract
Protein language models (pLMs) encode information about protein sequences which enable downstream tasks such as structure prediction, but their internal representations are not well understood. Sparse autoencoders (SAEs) provide a promising tool to disentangle latent pLM representations into interpretable features. We introduce an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein $\text{C}_{\alpha}$ backbone. Across ESM-2 8M layers, an FDR-controlled discovery analysis shows that local geometry is significantly associated with many SAE features, with varying levels of predictive strength,expanding coverage beyond database and sequence-based methods. This residue level annotation of SAE features reveal substructure within known biological labels. Geometric patterns critical to protein folding are discovered. A significant portion of SAE features activate on unannotated metagenomic protein sequences enabling us to use our SAE annotations to better understand these sequences.
Chat is not available.
Successful Page Load