Inverting the Bellman Equation: From $Q$-Values to World Models
Alistair Letcher ⋅ Mattie Fellows ⋅ Alexander D. Goldie ⋅ Jonathan Richens ⋅ Jakob Foerster ⋅ Oliver Richardson
Abstract
Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: while the former learns an explicit model of the transition dynamics $P$, model-free agents typically estimate value functions tied to a specific policy and reward. In this paper, we challenge this dichotomy by proving that value-based agents trained on a sufficiently rich set of reward functions, e.g. using goal-conditioned RL, implicitly encode a unique and accurate world model. To extract this model in practice, we introduce $P$-learning; analogous to $Q$-learning, which approximates $Q^\star$ using samples from the environment, $P$-learning extracts an agent's model of the environment $P^\star$ by sampling from its $Q$-values, policies, and rewards, effectively inverting the Bellman equation. In deterministic MDPs, we prove that the true kernel $P$ can be recovered from an agent trained on a single generic goal for finite state spaces $\mathcal{S}$, and a finite number of Gaussian goals for continuous $\mathcal{S} \subseteq \mathbb{R}^d$, provided $Q$-values are accurate. In the stochastic setting, these conditions generalise to larger sets of goals depending on the reward function family. Even when our assumptions are violated, we empirically demonstrate that agents trained with a small number of sparse rewards encode accurate dynamics in (i) stochastic variants of FourRooms, (ii) MountainCar, and (iii) Reacher. This is further validated by training policies inside the extracted world model, for goals far beyond the training distribution, suggesting that goal-conditioned agents secretly contain implicit generalisation capabilities and providing a new lens into the connection between model-based, model-free, and goal-conditioned RL.
Chat is not available.
Successful Page Load