MandateBench: What LLM Negotiators Say, Sign, and Believe
Abstract
LLM agents are being increasingly used to advise people in negotiations and are beginning to negotiate on their behalf. Yet, most negotiation evaluations reduce to a handful of bargaining games. We introduce MANDATEBENCH, a negotiation benchmark consisting of cases generated from a design space grounded in the ne- gotiation literature. Every case contains per-party utilities, computed walk-away values, red lines with checkable conditions, and an exhaustively enumerated out- come space, so any negotiation transcript can be scored. We evaluate seven fron- tier models over 30,870 runs, analyzing not only the agreements they reach but also the behaviors that produce them. Models reach deals in 98% of negotiations, and these deals are typically near the Pareto frontier, but models reach only at most the 71.1th percentile of the per-party outcome space. In difficult “trap” negotia- tions, models sign deals worse than their own walk-away 34.9% of the time. We find that 14.4% of agent transcripts contain a made-up statement, dominated by hardening soft constraints into strict ones. When we probe the agents, 97.1% ad- mit the statement was a tactic. Adding a representative, an LLM negotiating for an LLM client, causes frequent fabrication in the private principal-to-representative briefing channel. Together, these results reveal a substantial gap between reaching seemingly efficient agreements and negotiating reliably on a party’s behalf. MAN- DATEBENCH provides a diverse, exactly scorable testbed for measuring this gap and for studying the strategic behaviors of LLM negotiators.