Is a Linear Probe Evidence of a Linear Representation?
Abstract
The linear representation hypothesis (LRH) holds that high-level concepts are encoded as linear directions in a model's feature space, and grounds much of mechanistic interpretability, from probing to sparse autoencoders to steering vectors. We argue the standard test for it is too lenient. A high linear-probe accuracy shows that a concept is linearly readable from the representation, not that it is one of the representation's privileged directions of variance, which is what the LRH actually claims. These two readings come apart in any high-dimensional, richly structured representation, and the gap matters. The structure that linear probing ignores is precisely where compositional generalization, out-of-distribution behavior, and the validity of post-hoc feature decomposition live. We propose three asymmetric diagnostics that operationalize the strong reading: reverse predictivity, spectral concentration, and out-of-distribution generalization. We validate them in a fully observed simulation where ground truth is known, then apply them to pretrained vision encoders on SVG-World, our deterministic generative benchmark with paired visual styles. Pretraining produces structure beyond what a forward probe can certify, but the linear-feature reading systematically overstates how much. Probing on its own is not enough to license the mechanistic claims that rest on it.