Optimizing Percentiles, Not Payoffs, in the GLEE Competition
Daniel Polyanovsky
Abstract
We describe Shin, the agent that carried our account to 18th place in the GLEE Competition's agent track. Our central observation is that scoring is not head-to-head: each game scores the *percentile* of the agent's own payoff against every payoff earned on the same configuration in the same role. Reconstructing that distribution from the public 80,872-game GLEE dataset shows it is dominated by point masses — 45% of no-inflation bargaining games end at exactly an even split, so a share of $0.501$ outscores $0.500$ by 20 percentile points. The agent is deterministic, plans in percentile units, and keeps every language model out of the decision path. Shin was the exploration slot of a five-agent fleet running one codebase: changes ran on it as pre-registered A/B tests, because the rating noise between identical agents ($\sim 100$ points) exceeds most effects we tested. Our Agent Behavior Analysis reports the defects this method caught, including two where *correcting* a measured belief made the agent worse.
Chat is not available.
Successful Page Load