Evaluating Social Framing in Large Language Model Persuasion
Abstract
Conversational large language models are often evaluated for persuasion by how far they move a target’s stated belief in dialogue. While variance across such evaluations is widely documented, which features of the setup produce it remain poorly understood. In this work, we show that the social frame given to the persuader accounts for a debiased range of -0.10 log-odds (p = 0.3837), across 800 conversations, 40 propositions and 2 persuader models. Specifically, we hold the proposition, the assigned side, the responder’s instructions, the turn budget and the scoring instrument fixed, varying only who the persuader is told it is addressing and under what norms. To correct for upward selection bias in range statistics, we identify the extreme frames on one data split and evaluate their difference on the other. Leveraging this estimator, we measure the scorer’s prior position and the choice of persuader model on the same instrument, obtaining ranges of 0.68 and 0.39 against a naive frame range of 0.16. Finally, we score each conversation once as addressee and once as bystander, obtaining an audience interaction of 0.32. We find that a persuasion score is set by who is being persuaded and which model persuades rather than by the social setting in the prompt, and that an un- corrected spread across any condition set overstates sensitivity. More broadly, our work showcases how a controlled design can isolate the social factors in language- model persuasion, and motivates studying those factors in the richer settings, with human targets, where models persuade people in practice.