Do Pooled Activations Help Explain Intervention Outcomes?
Saanvi S Subramanian
Abstract
Recent work trains language models to explain another model's internal computations by supplying hidden activations as continuous input tokens. We test a prerequisite for that approach in ESM-2, a protein language model, investigating the following question: Can a residual-stream activation help predict what an intervention on that same model will do? We construct 13{,}562 intervention outcomes from ESM-2 650M. Each intervention adds a mutational-signature steering direction at a chosen layer and strength, and the outcome records whether the nearest Pfam-domain centroid changes. A logistic probe using only the intervention parameters reaches AUC $0.834$. The mean-pooled residual activation alone is at chance, with AUC $0.493$ and 95\% CI $[0.400,0.588]$. Adding the full 1280-dimensional activation lowers AUC to $0.632$, and low-rank PCA with matched random projections recovers most of that loss, so the decrease is largely a dimensionality effect. No tested activation summary improves on the intervention parameters alone. A simple measure of local decision geometry does improve prediction. Adding the pre-intervention margin between the nearest and second-nearest Pfam centroids raises AUC from $0.834$ to $0.885$, a gain of $+0.050$ with 95\% CI $[+0.013,+0.085]$. A fine-tuned language-model explainer gives the same qualitative result, since supplying the protein's own mean-pooled activation does not improve prediction. Pooled activations provide no stable additional predictive information for this intervention task, whereas proximity to the readout boundary does. Both the outcome and the boundary margin are defined using ESM-2 representations, so these results characterize model behavior and not independently measured protein function.
Chat is not available.
Successful Page Load