When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls
Juli Huang
Abstract
Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a standard way to infer which components of a language model are causally responsible for a behavior. We show, using GPT-2 small as a case study, that this recipe is fragile in ways that are rarely checked. First, we find that a natural implementation of “zero out head $h$” (overwriting a channel slice of the attention block’s returned output) does not ablate a single head at all: the block’s output projection mixes contributions across heads into output channels before the block returns, so this intervention zeroes an arbitrary slice of the already-mixed residual-stream update. On a controlled induction-style copying task, this naive method is essentially uncorrelated with a corrected, pre-projection ablation (Pearson $r=0.057$) and selects a completely disjoint top-5 set of “important” heads (0/5 overlap). Second, using the corrected method, we show that accuracy cannot detect further degradation at a behavioral floor (0.0% on an off-target arithmetic task). Near a ceiling (99.5% on the controlled induction-style copying task), it can detect large degradations but may miss confidence changes that do not cross the decision boundary. Log-probability of the gold continuation stays graded in both regimes. Third, with a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the per-head effect ranking is stable across the two splits (Spearman $\rho=0.974$ over all 144 heads), and the top-5 discovery-selected heads’ mean held-out effect substantially exceeds both control distributions (Monte Carlo $p=0.001$). Evidence for task specificity is not robust on GPT-2: the primary run is suggestive ($p=0.078$), while a single pre-committed larger rerun gives $p=0.274$. On DistilGPT2, the intervention-semantic and matched-control findings replicate, and the arithmetic floor recurs. Copying accuracy is lower (90.0%), so GPT-2’s near-ceiling example does not replicate. A single-head ablation result is not self-certifying: it requires verifying the intervention’s projection semantics, a continuous metric away from floor/ceiling, and matched controls evaluated on held-out data.
Chat is not available.
Successful Page Load