Extracting and Steering Ensemble and Structural Properties in BioEmu
Abstract
Generative protein models have recently demonstrated strong performance in protein structure prediction tasks. In particular, the AlphaFold models are the standard for static structure prediction. Other models like BioEmu add additional functionality by generating protein conformations trained to sample thermally distributed conformations. However, we still lack a mechanistic understanding of how these models sample the distribution of structures. To improve our underlying understanding of these models, we train sparse autoencoders on BioEmu and ESM3 to identify interpretable features corresponding to structural and thermodynamic properties. In addition, we assess at what point in the model these properties are most highly represented. From this, we determine that the hidden activations can be used to predict desired ensemble properties from a single sampling trajectory. Further, we show that the hidden activations can be directly tuned to quantitatively steer generation towards a desired property. These results improve mechanistic understanding of protein models and open the door for targeted property steering of protein structure sampling models like BioEmu.