On the Convergence of First-Order Methods in Signaling Games
Abstract
We study first-order learning dynamics in Bayesian signaling games, a minimal extensive-form model of strategic interaction under asymmetric information. In these games, a Sender observes a private type and chooses a signal, after which a Receiver forms a posterior belief and selects an action. This Bayesian belief update makes the induced learning field fundamentally different from the affine game-gradient fields of normal-form games. Empirically, we find a sharp dichotomy: standard online learning algorithms reliably converge to strict Perfect Bayesian Equilibria (PE), but not to non-strict ones. We show that this behavior cannot be explained by the global geometric conditions commonly used to prove last-iterate convergence, such as monotonicity or the Minty variational inequality. Indeed, these conditions can fail even in binary signaling games. Instead, we identify the relevant local geometry. Our main result characterizes variational stability in finite signaling games: a PE is variationally stable if and only if it is strict. This yields local last-iterate convergence guarantees for online mirror descent and regularized dual averaging near every strict PE. We then study global average-iterate convergence and identify a class of signaling games in which coarse-trigger regret-minimizing dynamics converge to the unique PE. Finally, we complement the theory with experiments comparing projected gradient ascent, mirror-based methods, optimistic variants, and PPO across canonical signaling-game families. The results position signaling games as a compact testbed for understanding how Bayesian belief formation reshapes the convergence geometry of multi-agent learning.