PID-DiffRL: Safe Offline Reinforcement Learning as Noisy Delayed Feedback Control
Abstract
We study safe reinforcement learning in the fully offline setting, where adaptive constraint enforcement is substantially harder than in online RL. In online Lagrangian methods, constraint violations are measured from fresh on-policy rollouts, providing relatively direct feedback for updating the safety multiplier. In offline safe RL, however, the multiplier must be updated using learned cost critics evaluated on actions proposed by the current policy, which can be out of distribution. This makes the feedback biased by critic extrapolation, delayed by target-network updates, and noisy due to bootstrapped value learning. We propose PID-Diffusion RL (PID-DiffRL), a safe offline RL framework that treats adaptive penalty learning as a feedback-control problem under imperfect critic-estimated signals. Our method replaces the standard integral-only multiplier update with a proportional-integral-derivative (PID) update to improve responsiveness and reduce oscillatory constraint behavior, while conservative cost critics reduce optimistic safety estimation under distribution shift. We further use diffusion policies instead of unimodal Gaussian policies, enabling the actor to represent multimodal safe action distributions that commonly arise in offline datasets and fragmented feasible regions. Empirically, PID-DiffRL achieves strong reward-safety trade-offs on DSRL benchmarks, consistently reducing constraint violations to near zero while maintaining competitive reward performance.