Coarse-to-Fine Category Structure in Child's-View Model Representations
Abstract
Humans can categorize the same object at different taxonomic levels, from broad kinds to fine-grained classes, supporting generalizations that vary in scope and specificity. Classical connectionist models show how these distinctions can emerge in learned representations from patterns of properties shared across objects. Previous studies, however, have relied on deliberately constructed concept–property data or conventionally curated image collections. We study whether comparable organization emerges in self-supervised vision models trained on naturalistic child's-view video. To characterize this structure, we progressively reconstruct model representations using increasing numbers of principal components and measure when category distinctions at different levels become most apparent. Our results show that coarse-to-fine category structure emerges from child's-view learning, and that broad distinctions become reliable before finer distinctions during training. The learned representations align more closely with human perceptual judgments than with conceptual or contextual judgments. Grounding these representations in caregiver speech leads to modest improvements in conceptual and contextual alignment. These results suggest that child's-view-trained models provide a useful setting for studying how multilevel category organization emerges from visual experience and what information the learned geometry captures.