Measuring How AI Agents Persuade and Negotiate: Outcome-Grounded Process Evaluation of Strategic Dialogue
Abstract
AI systems increasingly act as social actors in people's decisions: they persuade, negotiate, and take actions on a person's behalf. Measuring how such a system bears on a decision is a precondition for studying whether it supports a person's agency or exploits them. Large language model (LLM) judges can score the behavior, but agreement with human raters, their standard validation, shows that the judge scores as a rater would, not that what it scores corresponds to the actual outcome. We ask what recorded outcomes alone can establish about such a judge, using one simple test: given a theory-grounded rubric, do an LLM judge's reply-level and trajectory-level scores predict the final decision, at each stage of the interaction? In negotiation, under a strong reasoning judge, our rubric's scores cut prediction error by 41.3\% in the final quartile of the interaction. A generic natural-language-generation quality rubric, scored by the same judge, cuts it by 1.8\%. The behaviors our rubric scores turn from productive to harmful as the interaction progresses, an effect that whole-conversation scoring averages away. Human raters applying the same rubric agree with a lower-cost judge as closely as they agree with each other, and their scores predict the decisions comparably. With only the rubric swapped, the instrument transfers to persuasion for charitable donation, where the outcome is a human's decision to donate, cutting prediction error by 33.7\% in the final quartile against 1.6\% for the generic rubric under that same judge. The instrument gives developers behavior-level feedback and researchers a measure of how an agent's behavior relates to the decisions people make.