Protein Structure Model Embeddings Encode Interpretable Signals of Multistate Behavior
Abstract
Many proteins assume multiple conformational states as a core part of their function. Protein structure models often fail to adequately generate all of these conformations, instead defaulting to a single structure. We show that, despite this limitation, structure models still encode interpretable information relevant to multistate behavior. Simple classifiers trained on frozen embeddings discriminate operational multistate labels in a homology-grouped benchmark. This signal appears to remain after targeted controls for measured covariates and dataset source. Sparse autoencoders reveal latent features associated with residues relevant to conformational mobility and to protein-level multistate labels. Local structures at residues activating these latents show structural properties previously associated with conformational mobility. Results were broadly consistent across ESMFold and BioEMU embeddings, suggesting that this information may extend across learned protein structure representations. Conditioning, fine-tuning, and steering mature single-structure models toward multistate generation are therefore promising future directions.