Decentralized Q-Learning in Markov Potential Games
Onur Ünlü ⋅ Eilyan Bitar ⋅ Francesca Parise
Abstract
We propose a fully decentralized $Q$-learning dynamic for infinite-horizon discounted Markov potential games in which agents observe the global state and their own realized payoffs, but do not know the reward functions, the transition probabilities, the potential function, or the policies and the actions of the other players. In the proposed dynamic each player maintains a local estimate of its continuation payoff for each state and action and updates its policy through a smoothed best response. Importantly, the smoothing parameter is adjusted by using an online estimate of the discounted state-visitation distribution, so that players explore more in rarely visited states while exploitation is encouraged in states that are frequently visited. We prove that the induced policy sequence converges almost surely to a set of approximate Nash equilibria of the Markov potential game, with an approximation guarantee that depends only on easily computable parameters. We illustrate the performance of the proposed learning dynamic via numerical experiments on a routing game and compare it with existing independent learning schemes based on persistent $\theta$-greedy exploration.
Chat is not available.
Successful Page Load