Conservatism Controllable Compositional Guidance for Offline Safe Reinforcement Learning
Abstract
Learning constraint-satisfying policies from offline data without risky online interaction is crucial for reliable deployment of reinforcement learning (RL) agents. Representative methods leverage Implicit Q-Learning and advantage-weighted regression to learn value functions and balance the reward–safety trade-off, demonstrating solid safety performance in safety-critical hard constraint scenarios. However, these methods typically couple reward and safety optimization during offline training and rely on a predefined unified factor to balance the reward–safety trade-off (conservatism). Which ignores the fact that different tasks, due to their distinct reward and cost functions, may require different degrees of conservatism to satisfy safety constraints, resulting in sub-optimal performance. To address this limitation, we propose Co3G, a compositional generative model with test-time controllable conservatism for offline safe RL. During training, Co3G decouples the optimization objectives by separately employing reward-optimal guidance and safety-optimal guidance to induce high-reward and high-safety actions. During deployment, Co3G combines the reward and safety objectives via compositional generation, and further provides three mechanisms—manual tuning, rejection sampling, and RL—to balance the reward and safety guidance scales during deployment, thereby achieving a superior reward–safety trade-off. Extensive experiments on the OSRL benchmark show that Co3G delivers strong safety performance under stringent safety constraints while retaining flexible control over conservatism, significantly outperforming previous hard-constraint and soft-constraint baselines.