Deal With It! Diagnosing Post-Training in Bargaining, from Fluency to Surplus Extraction
Abstract
Remember how GPT-3 was fluent in English but notoriously lacked symbolic reasoning ability? Whilst reaching expert level coding or theorem proving abilities, in multi-turn bargaining, modern LLMs disappoint to extract economic value. Yet, although the domain's game theoretic basis is ideal for theoretical reference points, most evaluations only compare LLMs with each other, resembling leaderboards. In this paper we introduce theoretical upper and lower bounds, playing against scripted opponents for sets of simulated episodes. Upper: the optimal reward achievable with full knowledge of the opponent's behavior and reservation value. Lower: the highest reward attainable for a policy that acts unconditional to the opponent's previous actions (open-loop adjacent). In our prepared zero-sum multi-turn bargaining environment, a learner's achievable maximum lies in [R_open-loop, R*_full-information] = [0.350, 0.543]. A policy undershooting the lower bound fails to convert opponent offers into economic surplus. The subsequent evaluation of a frontier model underperformed the lower bound while producing fluent counteroffers, corroborating that fluency does not need to imply surplus extraction. Finally, attempting to leverage the benchmarked environment for post-training, a Qwen3.5-4B policy failed to exploit rollouts efficiently on vanilla GRPO. Despite seemingly stable training (mean reward -0.43 to 0.35), by step 50 over 60% of sampled rollouts suffered advantage collapse, as reported for GRPO in recent adjacent RLVR research.