Comparing Sparsified MLPs and Transcoders: Many Transcoder Features Are Neurons in Disguise
Abstract
MLP neurons in transformer language models are widely understood to be polysemantic, motivating transcoders — sparse, learned approximations of MLP layers — as interpretable replacement models for mechanistic analysis. However, recent work shows that MLP neurons can be used to recover sparse, interpretable circuits, raising the question of whether transcoders are necessary for circuit analysis. We define neuron features as transcoder features whose weights are strongly aligned with individual MLP neurons, and show that they account for a substantial fraction of active features in late layers across 13 open-source transcoders. Only neuron features have a stable correspondence with something in the underlying MLP: other features vary more between training runs, do not match the MLP's Jacobian, and change less significantly from their random initialization. Our findings reframe MLP neurons as a stronger basis for interpretability than commonly assumed, and call for better methods of attributing the non-neuron structure that is used by the model.