The Baby Intuitions Benchmark 2: Evaluating the Origins of Social Intelligence in Humans and Machines
Abstract
From early in development, the ability to infer the hidden social and object goals driving others’ actions is key to human social learning. Can advanced artificial intelligence infer such goals? To test this questions, we present the Baby Intuitions Benchmark 2 (BIB2), a benchmark designed to probe social and object-goal reasoning from minimal action cues. We report preliminary results from 74 adults across the full six-task benchmark, validating expected patterns of human performance, together with prior results from 15-month-old infants tested on two of the benchmark’s tasks. We introduce three computational approaches for evaluation on BIB2: (1) a baseline video transformer model which succeeded on BIB2’s predecessor, BIB; (2) a variant of this model that incorporates representations from the Baby Zero-shot Visual World Model (BabyZWM), pretrained on BabyView, a naturalistic dataset of head-mounted video collected from children within the first five years after birth; and (3) Baby V-JEPA2, a predictive video model also pretrained on BabyView which we further fine-tune on BIB2’s background training set. Initial evaluations of the baseline model trained on BIB and BIB2 background data suggest stronger alignment with adults on object-directed than social reasoning tasks; evaluation of the BabyView-pretrained models is ongoing. By probing humans’ and models’ reasoning about affiliation and agency from minimal cues, BIB2 provides a principled framework for studying the foundations of social intelligence and for advancing human-like social AI.