DeFlow: Decoupling Behavior-Prior Modeling and Value Maximization for Offline Policy Extraction
Abstract
We present DeFlow, a decoupled offline RL framework for extracting high-value actions from a learned multi-step flow prior. Directly optimizing iterative generative policies typically requires backpropagation through ODE solvers, while shortcut policies trade off expressivity, inference cost, and policy-improvement objectives. DeFlow keeps the flow model as a behavior prior and trains a lightweight action-conditioned residual module for value improvement under an adaptive trust-region penalty. This design bypasses solver differentiation and encourages proximity to the learned prior without treating the trust region as exact data-support preservation. Empirically, DeFlow is competitive with recent generative offline RL baselines, with the clearest gains on selected multimodal manipulation and offline-to-online settings.