Developmental Visual Experience Scaffolds Grounded Concept Acquisition in Vision-Language Models
Abstract
While vision-language models now achieve impressive image–text alignment, they often rely on surface-level associations rather than grounded conceptual representations that generalize across instances, support abstraction, and align with human semantic structure. Inspired by human visual development, in which infants begin life with limited acuity and chromatic sensitivity that gradually mature, we ask whether developmentally structured perceptual input can serve as an inductive bias for grounded concept acquisition. We trained vision-language contrastive learning models under two regimens: a standard regimen using clear images throughout, and a biomimetic regimen in which initially blurred and grayscale inputs gradually transition to clear images. Although both models achieve comparable caption-level alignment, the biomimetic model exhibits substantially stronger alignment at basic-level and superordinate category levels. Moreover, sparse autoencoder analyses reveal more compatible latent codes across modalities, and learning trajectories exhibit a clear coarse-to-fine progression, with broad distinctions acquired earlier in training and finer distinctions emerging later. Biomimetic representations also better match taxonomic structure and human behavioral similarity judgments. Taken together, these findings suggest that early perceptual limitations are not merely developmental obstacles but serve as adaptive inductive biases that scaffold structured, hierarchical, and human-aligned multimodal concept learning.