Do SAEs find steps in model processing?
Abstract
Circuits describe model processing as a sequence of ordered steps. For a feature to be part of a circuit, it must be a step, yet this important criterion is not explicitly verified. We propose punctuality as a metric of how concentrated a feature's emergence is to the layer where it was found, where highly punctual features arrive abruptly at the source layer and thus identify discrete steps of model processing. We find that sparse autoencoder (SAE) features are on average not very punctual. Crosscoders do find more punctual features than SAEs in later layers, but not more punctual than baselines. In the Bias-in-Bios SAE feature circuit, many features appear to be slicing points in time in a smooth, gradual refinement process, though there also exist punctual features which respect the orderability implied by their edges. Finally we explain that we should not unilaterally seek the most punctual features because punctuality turns out to be inversely correlated with feature persistence in later layers and magnitude of causal effect, both of which are desirable properties.