Plausible but Not Valid: Why Negotiation Evaluation Needs Counterparties That Respond
Abstract
Agents are increasingly negotiating on behalf of their users, and negotiation is relational i.e., how well it goes depends on what the counterpart does next. Yet these agents are evaluated against scripted counterparties rather than adaptive peers they encounter in reality. We run both arms in a single six-player bargaining environment, holding everything but this opponent class fixed, over 91 games and 118,758 events. The scripted arm passes its standard sanity check while measuring different consequences. It supplies a seventh of the negotiation per agent-seat, 9.2 inbound offers against 62.2, and never counters: zero across 1,531 bot proposals, with no bot having a counter path. It also changes which move the agent uses its turn on, opening more and countering less against the fixture, but the total per-episode count looks balanced, hiding this reallocation. It provides no response gradient either: at a fixed one-card request, four of five seats have an accept rule that ignores the offer, verified at 1,431 of 1,431 bot decisions. Yet the reasoning looks the same in both arms; the opponent-class signal is five to eight times weaker than the same classifier's read on authorship. Only the consequences diverge. Compliance with an explicit rule stated in prompt falls from 81.3% to 70.4% (OR 2.15; the scripted pool is also weaker, and a difficulty adjustment absorbs 8.2% of this), and a refusal shifts from an inventory miss (55.4%) to a price judgment (4.2%). Behavioral similarity is therefore not sufficient evidence of validity: it is the quantity a fixture reproduces best. Consequence-level claims about negotiation need counterparties that generate consequences.