Observer Effects in Mechanistic Interpretability
Abstract
Interpretability methods intervene on a model to read out its mechanisms, but the intervention can itself change the model's natural state: an activation forced off the data manifold may drive a pathway that is silent in the ordinary forward pass, fabricating a circuit that is an artifact of measurement rather than a feature of the model. Such off-manifold divergence is pervasive, and prior work shows it can manufacture illusory attributions---yet not all divergence is consequential, and which divergences are \emph{pernicious} rather than \emph{harmless} has remained uncharacterised. We study this on the polysemantic SIIT benchmark, transformers with known, overlapping task circuits where a false positive is well defined, and propose a two-axis mechanistic typology of dormancy---\emph{gate-crossing}, where a component's own patch crosses a nonlinear boundary, and \emph{coupling}, where it displaces the live circuit off-manifold through interaction---with an instrument for each axis: a Curvature Score and an interaction-weighted off-manifold ripple. Across intervention regimes these signals separate pernicious from harmless components precisely where off-manifold magnitude cannot, and the two axes prove complementary. We present this as an exploratory measurement framework: falsifiable mechanistic signatures for when an intervention pollutes the very computation it aims to reveal.