Generating Critic-Guided Synthetic Experience in Latent World Models
Jorge de Freitas ⋅ Christos Ziakas ⋅ Alessandra Russo
Abstract
While offline reinforcement learning can extract capable policies from static datasets, pushing performance beyond the behavior distribution requires interactive experience that is potentially unsafe or limited in real-world applications. World models offer an appealing alternative: a synthetic sandbox where an offline agent can safely explore and improve before deployment. In high-dimensional visual manipulation, however, training inside a learned latent simulator can collapse as compounding model errors corrupt policy and value learning. We propose using the world model as a local transition generator, leveraging critic guidance to choose actions inside that model to generate synthetic experiences. We study principled action refinement methodologies, including Flow Map $Q$-Guidance (FMQ), a trust-region update along critic gradients. On OGBench robotic manipulation tasks, FMQ-guided synthetic transitions improve a frozen-latent offline policy by a relative improvement of 10.8\% on the average success rate, requiring no real-environment interaction and no test-time planner. These results suggest that world models can support offline reinforcement learning when synthetic interaction is local, task-grounded, and prevented from corrupting value estimates.
Chat is not available.
Successful Page Load