Deriving Causal Subspaces in Transformers
Andrew Jun Lee ⋅ Selma Mazioud ⋅ Kyle J Ray ⋅ Jasmina Urdshals ⋅ Paul Riechers ⋅ Adam Shai
Abstract
To what extent can causal subspaces be derived a priori from the transformer architecture? For each hook between a transformer's final MLP and its next-token predictions (i.e., the final layer's $\mathtt{resid\_post}$, $\mathtt{resid\_final}$, and the logits), we derive a causal subspace and deduce the representation it contains by first deriving the causal subspace of the logits based on the shift-invariant property of softmax, and then solving for the contents of this subspace in terms of each prior hook. In doing so, we find that the causal structure of activation space is not always captured by a single fixed linear subspace. Furthermore, we find that a causal subspace of the final MLP's input (and of every prior hook) can span the entire activation space because of the MLP's nonlinear function, though a smaller causal subspace can still be estimated, to first order. We provide empirical evidence for the derived causal subspaces and the estimated causal subspace of the final MLP's input, as well as preliminary findings on the function of the final MLP.
Chat is not available.
Successful Page Load