Iterative Latent Refinement for Value Learning in Offline Goal-Conditioned RL
Daoxin Li ⋅ Songcheng Xu ⋅ Guozhang Chen
Abstract
Offline goal-conditioned reinforcement learning (GCRL) enables goal-reaching from static datasets, but the learned value, critic, or compatibility score must support both long-horizon reachability inference and offline policy extraction. We study iterative latent refinement as an architecture-level replacement for the feedforward value or critic backbone used by existing offline GCRL algorithms. Instead of emitting a score after a single feedforward computation, the network repeatedly updates a latent representation using shared recurrent update weights and lightweight step-specific parameters before exposing the final actor-facing signal. On the recent benchmark datasets reported in this study, recurrent variants improve performance across several algorithms without changing their losses, data, actors, or training protocol; across the 21 reported algorithm-dataset rows, the mean improvement is $10.7$ percentage points, with the largest gains on stitching and bottleneck maze tasks. CRL diagnostics are consistent with an improved policy-extraction interface: recurrent critics increase normalized separation between matched and mismatched goals and produce actors that remain closer to dataset actions under several behavior-support proxies. Depth and update-capacity studies further suggest that the effect depends on how computation is allocated between recurrent depth and per-step expressivity.
Chat is not available.
Successful Page Load