AI-Generated Proteins as a Test of Sparse Feature Generalization in Protein Language Models
Abstract
AI-based protein design methods enable generation of large, diverse sets of artificial variants, some of which retain biologically relevant properties despite substantial sequence divergence from natural proteins. Recognizing such designs could support candidate prioritization for experimental testing, directed evolution, and more robust biosecurity screening. Protein language models (PLMs) offer a promising approach, as their representations encode abstract information about protein sequence, structure, and function, and interpreting these embeddings using methods such as sparse autoencoders (SAEs) could reveal the protein properties supporting recognition. Here we apply an ensemble of more than 150,000 designs generated across six diverse protein targets using two different design methods to investigate how SAE features behave across artificial sequence variation, whether they expose localizable and interpretable states, and how both PLMs and SAEs compare with established sequence-based methods for recognizing distant structural analogs. Wild-type-associated sparse features recurred across structural analogs despite substantial sequence divergence, including a residue-localized state that tracked predicted local structure rather than sequence identity. Sparse profiles largely preserved the signal already present in the underlying dense ESM-C embeddings: retrieval performance varied across targets and SAEs did not consistently exceed dense embeddings or profile HMMs. The SAE’s contribution therefore lies in decomposability rather than improved ranking, exposing localized feature states that can generate hypotheses about local biophysical constraints.