When Offline Proxies Fail: Evidence-Gated Policy Iteration in the GLEE Competition
Abstract
Offline replay is attractive for developing strategic agents, but its value is uncertain when evaluation depends on a changing population of reactive opponents. We study this problem through mosskappa-baseline-v1, an entry in the GLEE Competition spanning bargaining, negotiation, and persuasion. Development used held-out replay, bounded clean-agent screens, and decision-level ledgers; the submitted runtime was deterministic and made no external model calls. A negotiation intervention transferred from replay to a clean screen, whereas a high-AUC response model and several promising persuasion slices did not yield a deployable policy. During the main intervention window, official Overall rose from 1571.75 to 1850.46 and rank moved from 57 to 51, with sharply different movement across the three game families. The evidence isolates three reasons offline gains failed to transfer: the target behavior was too rare, a strong slice did not compose with the rest of the policy, or the public proxy did not transport to the private opponent pool. These failures motivate a practical promotion rule for interactive agents: establish exposure, policy-level value, and live transport separately.