A Primal-dual Approach for Semi-Infinitely Constrained Reinforcement Learning
Abstract
We propose a primal-dual policy optimization method for reinforcement learning in semi-infinitely constrained Markov decision processes (SICMDP), which extends standard constrained Markov decision processes by allowing a continuum of constraints. We call our proposed method \textbf{S}emi-\textbf{I}nfinitely \textbf{P}rimal-\textbf{D}ual \textbf{P}olicy \textbf{O}ptimization (SI-PDPO). By introducing an infinite-dimensional Lagrange multiplier, we derive the saddle-point formulation of the original SICMDP problem and prove strong duality under mild conditions. We then introduce a regularization term for the dual variable, and solve the resulting regularized saddle-point problem using first-order methods with a policy optimization subroutine (for example, NPG or PPO). Unlike existing primal-type policy optimization algorithms for semi-infinitely constrained reinforcement learning, our method does not require solving an inner-loop optimization problem, thereby mitigating computational intractability and reducing the over-conservative bias. We perform a series of numerical experiments, including a real-world power control task in wireless communications, to evaluate the performance of SI-PDPO. Empirical results show that SI-PDPO consistently outperforms existing primal-type algorithms.