Personalized and Collaborative Online LQR via Thompson Sampling
Shivam Bajaj ⋅ Prateek Jaiswal ⋅ Vijay Gupta
Abstract
Reinforcement learning (RL) is well known to be data-intensive, and recent works have proposed leveraging data from similar systems to improve sample efficiency. In this work, we focus on the collaborative Linear Quadratic Regulator (LQR) setting and study the problem of learning under heterogeneous system dynamics. Due to the presence of a heterogeneity-induced additive bias, most existing collaborative methods are suitable only for low heterogeneity regimes and exhibit suboptimal performance under large heterogeneity. Based on the principle of Thompson sampling (TS), we propose an algorithm that, by leveraging data from other agents, yields a \emph{personalized} controller for every agent without incurring any additive heterogeneity-induced bias, i.e., with sublinear Bayes regret even under large heterogeneity. Under high heterogeneity regimes, we establish that our algorithm incurs $\tilde{\mathcal{O}}( T^{1-0.5\epsilon}\sqrt{\Delta})$ cumulative regret, where $T$ denotes the time horizon, $\epsilon\in[0,1]$ is a user-defined personalization parameter, and $\Delta$ denotes an upper bound on the maximum dissimilarity among the agents' dynamics. Our algorithm and analysis also generalize the special case of no heterogeneity. In particular, under no heterogeneity, we establish that our algorithm incurs $\mathcal{O}(\sqrt{T/M})$ cumulative regret. From a distributed implementation perspective, our method incurs only logarithmic communication overhead. Additionally, our algorithm outperforms methods that do not personalize data from other agents and, in certain regimes, also outperforms methods that do not utilize any data from other agents.
Chat is not available.
Successful Page Load