Representational Structure of ConceptARC Transformations in Pretrained Vision Models
Abstract
Abstract visual reasoning depends both on the structure of visual representations and the reasoning processes that operate over those representations. In the present work, we asked whether pretrained visual models contain structure corresponding to human-defined transformation families in ConceptARC, a benchmark designed to test visual concepts involving objects, space, geometry, and number. Across 160 tasks from 16 ConceptARC families, we represented each input-output transformation as the direction between its embeddings in two frozen vision models, CLIP and DINOv2. In both models, transformation directions were more similar for tasks belonging to the same concept family than for tasks from different families. Moreover, this effect exceeded the effect of decoding from static input and output representations alone, and from a rich hand-engineered representation of visual change. Concept-level organization was also highly similar across encoders, suggesting that the encoders represented concepts similarly. These findings suggest that general-purpose visual learning can produce reliable representational structure corresponding to conceptual transformations. More broadly, these findings support the possibility that part of what appears to require reasoning in visual abstraction tasks may already be supplied by generic visual processing.