Tracing Model Influence Through Agent Decisions
Abstract
We measure a learned component's behavioral influence along the agent's executed computation: predictions, candidate values, planner selection, serving transformations, and returned actions. In the 2026 GLEE competition, pooled learning raises acceptance AUC from 0.400 to 0.853 on 5,123 game-held-out complete-information, unknown-horizon responses. A separate game-disjoint replacement experiment changes 57 of 128 returned actions at a nominal eight-offer budget and 76 at a 64-offer budget, with reconstructed initial beliefs held fixed. Finer candidate resolution raises the switch rate by 14.8 percentage points; valuing feasible output prices changes ten archived-model actions. Graded replacement yields increasing switch rates; retaining member-specific logit differences changes 72/128 actions. At the seller's reservation value, acceptance shifts from 11/12 in the training corpus to 19/81 in live play at a boundary represented in every model member. These interventions connect predictive improvements to the computations that determine agent behavior.