Exploration via Exploitation: The Blessing of Reward Diversity in Personalized Federated RL
Abstract
While federated reinforcement learning has traditionally guaranteed high efficiency in homogeneous environments, collaboration is often hindered when agents possess distinct environments and goals. In this paper, we flip this paradigm, framing reward heterogeneity not as a hurdle to consensus, but as a structural asset that catalyzes exploration. We first propose Personalized Federated Upper Confidence Bound Value Iteration (PF-UCBVI), which achieves optimal linear speedup by decoupling shared dynamics from personalized objectives. To eliminate the risks of explicit exploration, we then introduce Personalized Federated Exploration-Free Value Iteration (PF-EFVI), a purely greedy algorithm that leverages reward diversity to ensure state-action coverage. We prove that PF-EFVI attains logarithmic regret without explicit exploration bonuses under sufficient reward diversity. Our results show that agent disagreement is a vital resource that shifts the exploration burden from temporal complexity to spatial diversity, enabling safe and efficient collective learning.