A Reproduction Study of Weight-Based Mechanistic Interpretability in Bilinear MLPs
Abstract
Mechanistic interpretability typically explains a trained network by analyzing its activations, at the cost of training auxiliary models such as sparse autoencoders (SAEs). Pearce et al. (2025) pursue an alternative: bilinear MLPs, whose omission of element-wise nonlinearities makes each layer an exact quadratic form, so interpretable features can be read directly from the weights via eigendecomposition of the layer's interaction tensor. We present a systematic reproduction of that work. The vision claims reproduce fully: regularization induces low-rank, visually interpretable eigenstructure, and our ablations sharpen the original account by identifying weight decay, rather than noise augmentation, as the dominant cause. The language claims reproduce only partially: we confirm the qualitative discovery of sentiment-negation circuits, but the reported prevalence of low-rank feature interactions holds for only one of three publicly released models (a second becomes consistent once rarely-active dictionary features are excluded). We identify two candidate factors that the original paper leaves unreported – SAE training duration and the underlying model's training compute – and a third that is structural: the public unavailability of matching SAE artifacts, which makes part of the original configuration impossible to replicate exactly. Beyond reproduction, we test whether weight-based features are genuinely structural rather than dataset-specific: regularized bilinear MLPs transfer across handwritten-digit datasets and recognize geometrically similar letters, and we introduce Quadratic Form Similarity, a weight-space metric that separates structurally similar from dissimilar class pairs where eigenvector cosine similarity cannot. Finally, we show on MNIST-scale models that the low-rank structure the original paper discovers post-hoc can instead be enforced during training via Canonical Polyadic (CP) decomposition: at matched full training, CP models match dense accuracy and effective rank, and a bridge experiment connecting the two extensions shows their decision surfaces substantially coincide with the dense ones – enforced and discovered structure converge. Our results are consistent with weight-based interpretability as a viable paradigm, while demonstrating that its reproducibility hinges on artifact availability and training-compute details that interpretability research rarely reports.