Early Visual Degradation Facilitates Few-Shot Word Learning from a Child's Egocentric Input
Abstract
Children learn words from continuous egocentric experience while their visual system is still developing. Recent work has demonstrated that neural network models can learn word—referent mappings from longitudinal head-mounted camera recordings of a single child, but it assumed adult-like visual input quality. Here, we show that incorporating developmentally-motivated visual degradations facilitates word grounding from naturalistic egocentric input. Specifically, we trained models on either full-fidelity inputs throughout or on initially blurred and color-depleted inputs during early training phases. Our results reveal that the developmentally inspired model learns more compact visual representations and outperforms the standard one on a novel word-learning task based on few-shot visual examples. The benefits of early visual degradations are particularly pronounced for basic-level categories and strengthen shape biases akin to those observed in infant learning. These findings suggest that early visual degradations can provide an inductive bias for sample-efficient grounded word learning.