Auditing Frozen SAE Readouts as Evaluation Protocols: Spatial Extent Exposes Readout Fragility
Abstract
Sparse autoencoders (SAEs) are increasingly used to interpret vision-model representations by mapping small sets of latent features to concept-level scores. Such a frozen readout functions as an evaluation protocol: it is used to support claims about whether concept information is present, yet its validity under concept-preserving variation is rarely stress-tested. We audit a released vision SAE on a frozen CLIP ViT-B/32, using COCO instance masks as external ground truth for concept presence and spatial extent. On held-out images, frozen-readout reliability declines sharply as objects become smaller, with binary detection rising from approximately 28% at the smallest extents to 93% at the largest. Controlled same-object rescaling, with identity and background fixed, establishes a graded causal effect. To distinguish evaluation-protocol failure from degradation that remains under trained linear readout, we compare the frozen SAE readout with trained linear probes on the raw CLIP representation and full SAE latent. Spatial-extent-linked degradation persists under trained readout, but the frozen readout loses 0.259 AUROC from large to small objects versus approximately 0.12 for the trained probes, yielding 2.1× the degradation of the three-probe mean. We also test this on a second SAE from the same family and observe the same pattern. Thus, a frozen interpretation readout can become unreliable while substantial concept information remains linearly decodable from the underlying representation. Our results motivate treating fixed interpretation readouts as evaluation protocols whose robustness and inferential validity should themselves be audited.