LASER: Latent Space Adjoint Matching for Support Constrained Entropy Regularized Offline RL
Abstract
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent generative-model-based approaches mitigate this by performing RL within a constrained latent space, but they often sacrifice policy expressiveness or suffer from vanishing gradients. In this work, we identify that entropy regularization is essential to prevent mode collapse in latent-space RL, though applying it without introducing numerical instability or compromising expressiveness remains a significant open challenge. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized RL with expressive flow policies. By design, our approach satisfies dataset support constraints and prevents mode collapse. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses a single, constant set of hyperparameters to outperform baselines even with their hand-tuned hyperparameters, highlighting its robust applicability.