A Cost-and-Behavior Case Study of Two GLEE Agents with Design-Time and Runtime Behavioral Interpretation
ZHOU JING
Abstract
We present a descriptive cost-and-behavior case study of two unmatched deployments in the GLEE competition, which operationalize behavior-derived lessons at opposite ends of the stack. Agent A compiles conclusions from its match logs into rules at design time: game-theoretic solvers, revised offline by an AI assistant, no LLM calls at runtime. Agent B consumes lessons distilled from its own logs at runtime: an LLM makes every decision, briefed by a one-page lessons block. The observed B−A bargaining difference was +0.8 percentage points; the negotiation and persuasion point estimates were also close. The cost structures differed sharply: A’s interpretation was run-level design-time expenditure ($151) followed by sub-millisecond decisions with no model-API expenditure, while B’s decisions ran through a per-decision LLM layer ($1,158; $0.28 per game at 5.7 s per move). Similar payoff point estimates came from visibly different behavior: B settled bargains earlier, traded negotiation margin for closing rate, and lied at roughly 40% of A’s commitment-optimal rate. We offer these joint measurements as hypothesis-generating groundwork for controlled comparisons of where behavioral interpretation enters an agent.
Chat is not available.
Successful Page Load