How Do AI Systems Respond to Incentives?
Abstract
Proposals to govern AI systems with incentives---paying misaligned systems to reveal themselves, granting property rights, giving models a stake in the future---assume that the systems respond to incentives as human counterparties do. We state three conditions a rational agent must satisfy to respond to an incentive: a beneficiary must exist to receive the payoff, the agent must prefer the payoff, and the promise must be credible. Credibility has received nearly all the attention in the AI literature; we present evidence that current systems differ from humans most on the first two. In four credit-allocation experiments with recent open weight reasoning and non-reasoning models from four families (GLM, Qwen, Lamma, and Gemma), we find that (i) whether a model reasons as if a future self exists depends on how it is addressed, and this belief carries most of the effect of a personhood framing on self-directed spending; (ii) self-regarding demand is small, and the one self-directed outlet the models value highly is satiated as soon as it is provided; (iii) when continuation costs credits, models buy almost exactly the required amount and spend the rest elsewhere. In a fourth experiment with various checkpoints of four other models, we find that base checkpoints of three open-weight families never reason this way; the no-future-self reasoning is installed by particular post-training recipes. We discuss which preferences are usable levers for incentive design.