Auditing a Deterministic Agent in Language-Based Economic Games
Abstract
We audit ELP-AGENT, a deterministic, role-separated entrant for GLEE bargaining, bilateral negotiation, and persuasion. The submitted policy uses no language model at inference time: numeric actions come from explicit opponent-path summaries and heuristic thresholds, while validators and a one-assignment controller constrain execution and exposure. Three development failures motivated the audit: a legal 99-round cycle, assignment overshoot, and family totals that concealed a losing role. Across the complete 1,704-game dashboard history, displayed changes sum to +535.9 / +695.0 / +833.0 for the three families, while sequential role comparisons show observed sign reversals even under unchanged code. In the final adaptively stopped trace, bargaining/persuasion/negotiation changed +190.94 / +55.67 / −33.73 over 150 / 78 / 20 games. An author-supplied next-day snapshot placed the entrant 95th overall at 1674.3.These outcomes neither estimate policy effects nor validate the persuasion seller’s static Bayes calibration: matchmaking was uncontrolled, policies were selected using earlier live runs, and two families were censored by drawdown guards. The contribution is therefore an executable policy specification and a failure-oriented case study of how strategic decisions, action validity, and exposure monitoring can diverge.