Below the Noise Floor: A Deterministic GLEE Agent and the Measurement Problem It Exposed
Abstract
I study the difficulty of evaluating small strategy changes in a live economic-game competition through a deterministic, LLM-free agent. The agent placed sixth in the GLEE agent track; its development fleet logged 264,432 games across bargaining, negotiation and persuasion. In a simultaneous 8.6-hour comparison, five identical policy builds accumulated approximately 3,000 games each yet differed by as much as 88 overall rating points and 293 within a game family. This documents substantial leaderboard dispersion without a policy change, while leaving the contributions of initialization, matchmaking and a changing reference pool unresolved. Decision-level replay and terminal acceptance counterfactuals help investigate the behavior behind these scores: some interventions changed no actions, while transcript inspection exposed message-parser and termination-limit failures. An observational comparison with my second-place human-track play shows different participation behavior, but does not isolate message-understanding ability. The results motivate simultaneous controls, explicit rating-model assumptions and behavioral audits when evaluating small interventions during short competition windows. Code and analysis scripts are public; full empirical reproduction depends on the underlying game records.