Do Large Language Models with Cultural Personas Behave Like Humans?
Abstract
Large language models (LLMs) are increasingly used in social simulation experiments to make claims about human populations and to predict events in the real world. However, human behavior is culturally situated, so reliable social simulators need to faithfully model cultural variation. Existing cultural benchmarks score what a model says rather than what it does, so these scores do not reflect whether it would behave like the population it simulates, also known as the "Value-Action Gap". To address this gap, we instantiate three economic games from published human studies that each established a distinct cross-cultural behavioral pattern, antisocial punishment, national parochialism, and shared stereotypes, and together provide a broad geographic coverage spanning 44 countries. The games are Public Goods Game with Punishment (PGGP), Parochial Cooperation Game (PCG), and Cooperative Stereotypes Game (CSG). We evaluate the cultural fidelity of 39 open-weight LLMs on all three games with cultural personas across model families and sizes, for 19,600 rendered personas in total. We discover that while cultural personas do move behavior, fidelity is varied and limited, with scores reaching as low as 0.00 and best-in-game scores of only 0.78 (PGGP), 0.74 (PCG), and 0.72 (CSG). Fidelity does not transfer reliably across games, with a mean pair-wise Spearman correlation of only +0.14 and none significant, and it does not grow reliably with model size. Distributional behavior, on the other hand, does transfer under the same test on the same models, as a model that behaves like the "average" persona on one game tends to do so on the others, and fidelity weakens as populations grow more culturally distant from the United States. Our findings suggest that cultural conditioning alone does not make a model a suitable replacement for human participants in social simulations, and that their fidelity requires closer scrutiny before such use.