SMORE: Minimal-Pair Videos for Social, Moral and Rational Reasoning in MLLMs
Abstract
Understanding human behavior goes beyond recognizing objects and actions. It requires reasoning about why people act, whether their actions are rational, whether they help or harm others, and whether their behavior is morally appropriate. Importantly, these judgments can change when only a small part of an interaction changes while the broader context remains the same. It remains unclear whether multimodal large language models (MLLMs) can reliably reason about such subtle behavioral differences. We introduce SMORE, a minimal-pair video benchmark for Social, MOral, and Rational reasoning: 427 pairs (854 videos) and 1,281 questions in which paired videos share the same actors, environment, and overall interaction but differ in a single socially relevant factor. We evaluate models in two settings: paired-video, where both videos in a pair are shown together, and single-video, where each video is shown independently. From the latter, we also report Joint-Pair Accuracy, which requires both videos in a pair to be answered correctly. Humans achieve 96% Joint-Pair Accuracy, while the best model, Gemini-3.5-Flash, reaches only 62.4%, revealing a substantial gap in reasoning about the behavioral differences underlying social, moral, and rational judgments.