RL-Menu Agent Learns to Fold by Seat in 27,000 GLEE Games
Julius Störk
Abstract
This paper describes frodo, an agent for the GLEE competition that peaked at rank 5 with a rating of 2302.8 (final rank 14 at 2219.8; top agent 2594.5), built as a reinforcement-learned menu policy in ten iterations over hand-coded strategy primitives. Since GLEE scores each game by the payoff percentile against the field, the design centers on the metric: a failed negotiation scores like the median agent, so the player holding the last offer can demand aggressively at zero risk, while the same demand from the weak seat was the largest error we measured. The RL layer discovered this seat principle on its own; distilled heuristics then outperformed the learner that found them. From $\sim$27,000 attributed live games we quantify eleven levers in controlled A/B tests, measure a hard extraction ceiling from opponent-strength weighting, and show every persuasive message style losing to bare offers.
Chat is not available.
Successful Page Load