Generative Latent Priors for Interpreting and Steering Protein Language Models
Abstract
Sparse dictionaries are the dominant tool for interpreting Protein Language Models (PLMs). Generative Latent Priors (GLP) are a complementary approach: a diffusion model that learns the distribution of a frozen model's activations directly, imposing no assumptions. We port GLP to ESM-2 activations (UniRef50) and evaluate it at two host scales (ESM-2-8M and -650M), both as a lens for PLMs representations and as a prior for editing them. Generation quality scales smoothly with depth and compute. As a lens, single GLP "meta-neurons'' are the strongest probes of discrete Swiss-Prot concepts at both hosts, beating InterPLM's released SAE features, and lead on continuous properties at 8M though not at 650M; falsification tests show this is mostly concentration, not new information, though a handful of concepts resist every sparse dictionary combination. As an editing prior, the on-manifold correction makes diff-of-means steering of three biochemical properties effective and restores on-manifoldness, but only in isolation: re-embedding the sequences actually produced, the correction no longer helps and buys no structural plausibility. This is mainly reproduced in additional SAE-based clamping structural steering experiments, where only Helix motif succeeds. The naturalness failure is traced to the discrete decoder rather than the prior, highlighting what the method must clear before it can serve for both discovery-by-editing and discovery-by-exploration.