Sampling the Training Trajectory of Vision Models Reveals a Mismatch with Human Visual Development
Abstract
In this work, we compare time-resolved EEG representational geometry across seven age groups, from 6 months to adulthood, with representations sampled throughout training from diverse deep neural networks (DNNs). We analyze supervised and self-supervised ResNet50 and ViT-B/16 models. Previous model-brain comparisons mostly relied on adult brain recordings, focusing on untrained and fully trained networks. We find that successive training stages do not systematically progress from infant- to adult-like neural representations: the age group a model aligns with best is typically established early in training and remains stable thereafter or changes non-monotonically across training. Learned DNN features show their clearest advantage over their untrained counterparts for adult neural representations, while at younger ages their contribution is architecture-dependent, with ViT showing stronger learned-feature alignment than ResNet50. Finally, we show that neural representational content differs across development, with stronger sensitivity to low-level visual properties, particularly spatial frequency, at younger ages and greater categorical organization in adults.