Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction
Abstract
Enforcing state-wise safety constraints is critical for the application of reinforcement learning (RL) in real-world problems, such as autonomous driving and robot manipulation. However, existing safe RL methods often enforce state-wise constraints only in expectation, while hard state-wise guarantees typically require strong dynamics assumptions. The former can still allow rare but severe violations, while the latter is impractical in model-free stochastic systems. We instead target high-probability control of the maximum instantaneous cost along a trajectory. To accomplish this goal, we propose Absolute State-wise Constrained Policy Optimization (ASCPO), a model-free policy search algorithm whose exact update provides a distribution-free high-probability certificate for state-wise constraint satisfaction in stochastic systems. The guarantee holds under finite variance and does not require Gaussian, unimodal, or light-tailed cost distributions. We demonstrate the effectiveness of our approach by training neural network policies for extensive robot locomotion tasks, where the agent must adhere to various state-wise safety constraints. Our results show that ASCPO most consistently reduces state-wise safety violations while maintaining competitive reward across challenging continuous control tasks.