Before Words, Beyond Speech: Evaluating Nonverbal Social Reasoning in Early Childhood
Marie Amale Huynh ⋅ Laura Bravo-Sánchez ⋅ Lauren K Dubin ⋅ Nick Haber ⋅ Philip A Fisher ⋅ Serena Yeung-Levy
Abstract
Children's earliest years are a period of intense social learning, during which meaning is built through gaze, gesture, and physical contact. Yet, this critical aspect of social interaction remains under-evaluated in multimodal video understanding, as existing benchmarks prioritize the conversation-driven interactions of adults. We address this gap with NEST ($\textbf{N}$aturalistic $\textbf{E}$arly-childhood $\textbf{S}$ocial in$\textbf{T}$eractions), the first child-centric benchmark for nonverbal social reasoning. NEST comprises 1,208 manually annotated clips ($\sim$12s) curated for diversity across naturalistic settings, cultural contexts, and developmental stages. Beyond a lack of data, targeted model improvement in this domain is hindered by coarse evaluation schemas that conflate basic perception with higher-order reasoning. To resolve this, NEST introduces a compositional schema that builds atomic interaction cues into progressive levels of social reasoning: contextual, behavioral, interpersonal, and open-ended descriptions. Extensive analysis of state-of-the-art Vision-Language Models (VLMs) reveals a stark performance collapse: despite strong contextual recognition, models exhibit substantial gaps in decoding precise behavioral and interpersonal cues. We trace these failures to specific biases, including poor consequence grounding and temporal sensitivity. Finally, we ground NEST in real-world applications by evaluating VLMs using domain adaptation strategies. Ultimately, NEST provides a rigorous testbed for advancing socially grounded artificial intelligence in early childhood. The dataset and code will be publicly released.
Chat is not available.
Successful Page Load