Overfit on Purpose: A Rating-Safe Harness for Continual Adaptation in Live Economic Games
Liane Galanti
Abstract
We describe the agent we operated in the GLEE competition and the system around it. The agent itself is deliberately boring: a deterministic rule policy ($\sim$1{,}500 lines of Python, no model calls at decision time). The contribution is the \emph{harness} that continually rewrites it. An LLM operator mines the live competition's telemetry, distills observed opponent behavior into candidate rules, and deploys each change through a safety pipeline: a replay regression over 100K recorded decision states, a canary cohort on a credential-isolated clone agent, and an automatic \emph{bleed watchdog} that parks any (agent, game-family) pair whose rating drops past a threshold. The result is a continual harness for a competitive multi-agent field where every regression has an immediate, adversarially-priced cost. Our central empirical finding motivates the design: the \emph{same} rule earns rating at 1500 Elo and bleeds at 2400+, so rule value is opponent-level-dependent and no static policy survives. We report the loop, the safety mechanisms, three distilled-rule case studies (including a rule our own watchdog falsified twice in one day), and an analysis of how we verify the deployed agent does what the mined data says it should.
Chat is not available.
Successful Page Load