Facial-Impression Judgments of Multimodal Large Language Models Predict Election Outcomes
Abstract
Humans rapidly infer social traits of others from their faces, and these superficial judgments can predict consequential behavioral outcomes, including electoral choices. Given growing evidence that multimodal large language models (MLLMs) reproduce human-like facial impression judgments, we asked whether MLLMs' facial-impression judgments similarly predict real election outcomes. We presented MLLMs with paired facial photographs of the winner and runner-up in U.S. gubernatorial and Senate races held between 1995 and 2008 and asked which candidate appeared more competent. The models' competence judgments predicted the election outcomes in each race with 60--76\% accuracy, comparable to that achieved by aggregate human competence judgments. This predictive ability was consistently observed across 15 models spanning a wide range of general multimodal reasoning capabilities, as measured by the MMMU-Pro benchmark. There was no detectable relationship between election-prediction accuracy and general multimodal reasoning capability. Interestingly, when the same models were instead asked whom they would vote for or recommend voting for, their responses predicted election outcomes less accurately than their competence judgments. These results suggest that general-purpose MLLMs exhibit human-like facial biases that may appear innocuous yet are associated with consequential human behavior.