AdaptFlow: State-Anchored Goal Conditioning with Flow Matching for Offline Goal-Conditioned RL
Abstract
Offline goal-conditioned reinforcement learning (GCRL) aims to learn a single policy that adapts its behavior to arbitrary goals from a fixed dataset, making injection of the goal information into the policy a central design question. While recent works increasingly leverage expressive conditional generative modeling architectures, which have demonstrated remarkable success in computer vision, architectural choices for goal conditioning in offline GCRL remain under-explored. We identify that goals in GCRL differ from typical conditioning signals in two ways: they are often noisy or redundant, and they are consumed within a time-sequential episode where the relevant context of a fixed goal shifts as the agent's state evolves. Motivated by these observations, we propose State-Anchored Goal Conditioning (SAGC), a modulation-based goal conditioning scheme that derives its scale and shift parameters by jointly processing state and goal representations in a learned latent space. We further introduce AdaptFlow, an end-to-end trainable offline GCRL architecture that integrates conditional flow matching with SAGC. AdaptFlow outperforms matches all baselines on 23 out of 27 OGBench tasks.