Intuitive Physics in Video Models Trained on Children's Visual Experience
Abstract
Children demonstrate substantial knowledge about the physical world from an early age. Whether this knowledge can emerge from general-purpose learning algorithms when trained on data that reflects children's early experience remains an open question. We train three families of video models on BabyView, a dataset comprising more than 800 hours of egocentric video that captures the everyday experiences of young children. We evaluate these models' understanding of intuitive physics using the violation-of-expectation paradigm, and find that they can make meaningful predictions, particularly regarding object permanence, continuity, and constancy, while their understanding of other concepts, such as inertia and collision, remains limited. Interestingly, different model families exhibit distinct performance patterns across intuitive-physics categories, suggesting the importance of modeling design choices. Our findings indicate that general-purpose learning algorithms, as instantiated by video models, can support the acquisition of some physical concepts from children’s early visual experience. However, substantial limitations remain in both training and evaluation, and more work needs to be done to develop models with a robust understanding of the physical world.