The Intervention Lottery: Search-Matched Nulls for Reliable Discovery from Protein-Model Steering
Abstract
Activation interventions are often treated as stronger evidence than correlational feature interpretation: if steering an interpretable latent changes a relevant output and transfers to proteins held out from direction selection, the feature appears mechanistically validated. We show that this inference can fail. Across deep mutational scanning (DMS) stability landscapes, searching arbitrary hidden-state directions produces interventions whose mutation scores transfer strongly. We screen all 10,240 decoder directions of a public protein-language-model sparse autoencoder (SAE) and 1,024 norm-matched Gaussian directions on 50 source proteins, then evaluate 14 initial and 85 additional non-overlapping stability domains. Even after removing exact wild-type-to-mutant identity effects, source and external direction rankings correlate at ρ = .943 for Gaussian and .963 for SAE directions. The phenomenon persists from layers 2–6 and across ESM-2 models with 8M–150M parameters, but does not transfer to three GFP brightness landscapes. We formalize the intervention lottery curve: external performance expected after applying a specified selection rule to M null directions. A previously compelling hydropathy-feature intervention fails this stronger audit after mutation-type control. Causal output control and same-phenotype protein holdout are therefore insufficient evidence of biological mechanism; reliable discovery requires selection-matched nulls and genuinely shifted or prospective validation.