Sparse Autoencoder Features Causally Influence Active-Site Token Predictions
Abstract
Protein language models (PLMs) encode rich functional information, yet how this information is organized internally and whether specific features causally drive predictions remain unclear. We study enzyme-function representations in ESM-2 using layer-wise probing, sparse autoencoders (SAEs), and causal intervention. Functional information peaks at an intermediate layer rather than the final layer, suggesting that layer selection matters for downstream function prediction. A TopK SAE trained on this layer recovers EC-specific features that strongly localize to annotated active sites, with active-site residues ranking in the top 5\% of feature activations in 97.1\% of cases. Ablating these features selectively degrades active-site predictions, whereas matched random controls show no comparable effect. These results identify both where functional information is most accessible in ESM-2 and which sparse features are causally involved in its predictions.