Scale Is Not Color: Attribute-Dependent Limits of Causal Subspace Analysis in Vision Transformers
Javier Contreras ⋅ Ivan Sipiran
Abstract
Distributed alignment search (DAS) and related causal methods are now standard evidence that a model encodes a variable in an identifiable subspace. What licenses the opposite conclusion, when the intervention fails? We study a case where this question can be adjudicated. Recent work showed that large-scale pretrained Vision Transformers spontaneously implement a two-stage algorithm for abstract same--different reasoning: an early \emph{perceptual} stage building disentangled object-local representations, and a later \emph{relational} stage comparing them abstractly --- established on stimuli varying in shape and color. We replace color with object \emph{scale}, an attribute that is spatial and continuous rather than categorical, and rerun the full analysis pipeline (DAS, novel-representation generalization, linear probing with causal interventions) on six ViT-B backbones under two training regimes, with matched control conditions throughout: 96 DAS runs. Shape reproduces the published signature closely (peak intervention accuracy $\ge 0.90$). Scale never does ($\le 0.69$ in any configuration), and the two weaker methods agree. Critically, the same models solve the scale-based relational task behaviorally at 89\%, so the information is demonstrably present and used. The null is therefore evidence about representational \emph{format}, not about presence --- a distinction that a single method could not have drawn, and one that bears directly on how strongly causal-subspace results should be read. We argue that reporting which of these two claims a null licenses should be a routine standard for interventional interpretability.
Chat is not available.
Successful Page Load