A Constrained Bi-level Optimization Framework for Constrained Preference-Based Reinforcement Learning
Yue Mao ⋅ Siyuan Xu ⋅ Shicheng Liu ⋅ Minghui Zhu
Abstract
This paper studies the problem of jointly learning a reward function, a cost function, and a policy from an expert's preferences. We formulate the problem as a constrained bi-level optimization problem, where the upper level infers the reward and cost functions from preferences, while the lower level optimizes a policy to best align with those preferences. To solve this problem, we propose a double-loop algorithm, Constrained Bi-level Optimization for Preference-Based Reinforcement Learning (CB-PbRL), which solves the lower-level optimization problem in the inner loop and the upper-level optimization problem in the outer loop. We establish a theoretical guarantee that CB-PbRL converges at a rate of $\mathcal{O}(1/\sqrt{K})$, and we demonstrate its effectiveness across multiple simulation environments.
Chat is not available.
Successful Page Load