Payoff-Only Learning in Games via Best-Response Differential Inclusions
Ismail Hassan ⋅ Pedro G Lind ⋅ Anis Yazidi
Abstract
Each player in a finite normal-form game observes only its own realized reward; opponents, their actions, and their payoffs are unknown. We analyze an epochal payoff-only recursion in which the mixed strategy is held fixed across long epochs, own-action means are estimated from bandit data, and the iterate steps toward an empirical near-best vertex of a barrier simplex. Our main result reduces this stochastic recursion to the deterministic best-response differential inclusion (BRDI). Under a Hoeffding-tied epoch schedule, the empirical near-best correspondence is contained, on a tail event of probability at least $1-\sum_{k\ge K}\delta_k$, in a deterministic graph enlargement of the BRDI whose radius vanishes as $k\to\infty$. Without conditioning, the affine interpolation of the iterates is almost surely a bounded perturbed solution of the BRDI in the Bena\"im--Hofbauer--Sorin sense, and its $\omega$-limit set is internally chain transitive. Composing this reduction with deterministic BRDI confinement results yields last-iterate Nash convergence in three classes: $\operatorname{dist}(x_k,\NE)\to 0$ almost surely in finite zero-sum games, in finite exact-potential games (under an empty-interior condition on the potential's image), and in every generic $2\times 2$ game. Fixed barriers give $O(b)$-Nash equilibria of the original game; deterministic shrinking removes the bias.
Chat is not available.
Successful Page Load