GPT-4 reasons like infants about objects but not agents
Abstract
Theory of Mind (ToM) is central to human social cognition and is critical for successfully navigating the social world. However, most work to date has explored explicit ToM in AI, particularly with large language models (LLMs), and has found mixed results. In contrast, decades of developmental psychology research with human infants has found remarkable, early-emerging implicit ToM. However, little work has explored implicit ToM in LLMs. The current study investigates implicit social reasoning abilities in GPT-4 using model-facing versions of classic infant experiments. In Experiment 1, we explored GPT-4’s implicit ToM abilities across four social reasoning tasks adapted from infant paradigms. In Experiment 2, we explored whether GPT-4 was capable of physical reasoning to contrast with social reasoning in Experiments 1 and 2. Across these three experiments, our results suggest that there may be domain differences in GPT-4’s reasoning abilities, such that GPT-4 struggled with social reasoning tasks despite performing very well on physical reasoning tasks. These findings highlight the importance of systematic exploration of social reasoning in LLMs and provide insight into potential deficits in GPT-4’s social reasoning abilities.