Learning to Recommend in Unknown Games
Abstract
We study preference learning and coordination through recommendations in multi-agent game settings, where a moderator repeatedly interacts with agents whose utility functions are unknown. In each round, the moderator issues action recommendations and observes whether agents follow or deviate. We consider agents who best respond to the moderator's recommendations and study what this feedback reveals about their utilities. We characterize the class of games that are indistinguishable under this feedback model. Moreover, we introduce a notion of moderator regret based on agents' incentives to deviate from the recommendations and design an online algorithm with low regret under the best-response model, with guarantees that scale linearly in the game dimension and logarithmically in time. Our results lay a theoretical foundation for AI recommendation systems in strategic multi-agent environments, where recommendation compliance is shaped by strategic interaction.