Safe Offline Reinforcement Learning using Behavior Regularisation and Latent Feasibility-Guidance
Abstract
Offline safe reinforcement learning aims to maximize cumulative reward while satisfying safety constraints using only static datasets. Recent approaches leverage latent variable models to capture the behavior policy distribution, enabling efficient policy extraction under limited data. However, existing latent-space methods typically rely on soft constraints, which allow constraint violations in expectation and may be unsuitable for safety-critical settings. In addition, these methods often employ Advantage-Weighted Regression (AWR) for policy extraction, which exhibits exponential sensitivity to out-of-distribution critic bias, leading to high variance in the policy gradient estimate. To address these limitations, we propose Feasibility-Optimized Conditional Actor Learning (FOCAL). FOCAL enforces state-wise hard safety constraints by integrating reachability-based feasibility constraints directly into the Conditional Variational Autoencoder. To overcome the instability associated with AWR, we introduce a feasibility-partitioned, behavior-regularized objective that utilizes stable, first-order gradients to structurally decouple reward maximization in feasible regions from constraint violation minimization in infeasible regions. Finally, an adaptive test-time inference dynamically modulates latent sampling variance based on state-wise feasibility to ensure safe policy execution under out-of-distribution states. Theoretical analysis demonstrates that FOCAL circumvents the exponential variance in policy gradient estimates associated with AWR under critic errors, while providing formal bounds on policy deviation from the behavioral dataset. Empirically, on the DSRL benchmark, FOCAL satisfies safety constraints while achieving competitive or higher returns compared to state-of-the-art methods, all while maintaining efficient single-step inference.