What Can Activation Patching Actually Establish? Reachable-State Geometry and the Limits of Mechanistic Evidence
Ankush Kadu ⋅ Ananth Kalyanasundaram ⋅ Aswanth Krishnan
Abstract
Activation patching can yield stable, statistically convincing effects while remaining unable to distinguish the supported mechanism from plausible alternatives—a failure of the evaluation protocol, not sampling noise. We characterize the internal states that an intervention family $Q$ can reach in a trained network and, relative to a declared mechanism class $\mathcal{F}$, the blind distinctions that vanish everywhere the evaluation can observe. The algebra is classical aliasing. Our modelling contribution is that under patch-and-recompute the analyst chooses $q$ but the network determines $z(q)$, making resolving power measurable before any behavioural readout. On controlled operator slices, the $A \to B$ map determines the onset of exact polynomial blindness. In trained Transformers, a direction predicted prospectively from reachable-state geometry—without readout values or separation optimization—is practically invisible to single-site evidence yet visible to joint evidence in 172/185 held-out captures across three models and five behaviours. We expose support mismatch and same-state chart selection as two failure modes in protocol comparisons. On a published IOI circuit, single-site evidence detects non-additivity but poorly identifies the interaction: median allowable misattribution is 42.9%, versus 0.779% for matched joint evidence. Trustworthy mechanistic evaluation requires auditing not only whether an intervention produces an effect, but which competing mechanisms it could rule out.
Chat is not available.
Successful Page Load