Offline Evaluation under a Reconstructed GLEE Scoring System
Md Mustakin Alam ⋅ Aminul Islam
Abstract
The GLEE competition converts a payoff into a percentile against every payoff earned on the same configuration in the same role, seeded from a public dataset, so its scoring function is reconstructible offline. We built a deterministic analytic agent (closed-form backward induction against a calibrated opponent model, no language-model call at inference) and an evaluator that scores it against the released reference distribution before it plays anything. Over 12,840 live games in five agent slots, the evaluator ranked all four single-factor policy variants correctly and got the level badly wrong, over-predicting mean percentile by up to $0.107$. Splitting the error shows why: the simulator is close to unbiased ($|\Delta| \le 0.021$), while the released reference distribution is off by $+0.087$ in negotiation and, against our expectation, by $-0.021$ in persuasion. We also report an endgame defection that buyers partly price, and a randomised disclosure treatment the platform supplies for free.
Chat is not available.
Successful Page Load