The Question of Where: A Placement Asymmetry in How Agents Learn to Cooperate
Samridh Aggarwal
Abstract
A learning agent that holds a social preference must encode it somewhere between experience and action, and the reward is the obvious place to do so. We show that this choice of location matters more than the strength of the preference, and that what it sets is a property of the interaction: a single agent carrying the preference among unaffected partners is measurably indistinguishable from one carrying nothing at all. Using experience-weighted attraction learning in a two-player threshold public goods game, we compare fairness placed at action selection (choice-level), where a fixed fair prior guides the choice while the learned payoff signal is left intact, against fairness placed at the reward (learning-level), where a Fehr-Schmidt inequity transform rewrites the payoff before the agent learns from it. The same preference produces opposite outcomes. Reward-placed fairness drives coordination below that of a purely selfish learner, from 63.8% to 52.5%, while selection-placed fairness raises it to 72.4%. We prove that this is an asymmetry: reward placement can invert the cooperative ranking of a value-based learner, whereas selection placement provably cannot, for any rule that updates by weighted averaging. A closed form, $\omega^{*}=\rho/\Delta$, predicts the collapse before training. The divergence recurs across an eight-game battery, in a decoder-only transformer trained from scratch, whose coordination against a free-riding partner falls from sustained cooperation to near zero as the fairness weight crosses the predicted threshold, and in a pre-registered experiment with human participants. The result bears on how cooperative objectives are encoded in learning agents and reward models: a fidelity-preserving placement is a safe default, and placing a social goal in the reward shapes what the agent comes to believe, not just the incentives it faces.
Chat is not available.
Successful Page Load