Your Intervention Was Too Narrow: Mechanism Width as a Precondition for Interpretability-Driven Discovery
Sumit Yadav ⋅ Basanta Joshi
Abstract
Interpretability-driven discovery rests on interventions: a part is credited with a function when perturbing it changes behavior, and dismissed when nothing happens. We show that the field's standard intervention unit (one direction, one layer, one attention head) recovers only about 10% of the achievable causal effect in three unrelated testbeds, one task each: rank-1 steering in vision encoders (0.111), single-layer ablation in Llama-3.1-8B (0.096), and single-head sufficiency in GPT-2's IOI circuit (0.086). The mechanism's width $k^\ast$ can be measured before intervening, and widening the intervention to $k^\ast$ recovers 95-100%. The effect is about width itself: a 12-layer band does 7.5x the additive damage prediction, random subsets track contiguous bands, and pushing past $k^\ast$ destroys object identity. Across 12 language models, whether a single-unit test works is predicted by a cliff in the writer-score ladder (Spearman $\rho = 0.81$). Width is necessary, not sufficient: at a documented readout boundary, 81% of a representation gap closes while the consumer's output moves approximately 0, though it reads the attribute when prompted. A negative result at standard width is an underpowered test, not evidence of absence; a cheap protocol (sweep width, compare $k^\ast_{\mathrm{repr}}$ to $k^\ast_{\mathrm{eff}}$, run matched controls) tells the difference before a discovery claim is made.
Chat is not available.
Successful Page Load