WM-LAN: Test-Time Adaptation of World Models through Latent Refinement
Ankit Bhattarai ⋅ Matthew Macfarlane ⋅ Clem Bonnet ⋅ Florian Fischer ⋅ Per O Kristensson
Abstract
General agents that act in the real world need to adapt to environments they have never seen before. World models are a promising architecture for the core of a general agent, which it can use to plan, predict the consequences of its actions, and even learn. However, world models are only useful when they are accurate. When entering unseen environments, agents must continually adapt their world models on the fly in a sample-efficient manner. Amortised inference predicts using a single forward pass but is fixed once training ends, and degrades out-of-distribution (OOD) with no way to spend more compute to improve this estimate. In this work, we study the World-Model Latent Adaptation Network (WM-LAN), which adapts a world model at test time through latent refinement. WM-LAN uses an encoder--decoder architecture and then adapts by searching the latent space for the best latent representation that, when decoded, explains all data seen up to that point, requiring no reward, labels, or parameter updates. On CartPole and Walker with unseen gravity and actuator strength, WM-LAN removes $51$--$76$\% of OOD one-step error from just 32 observations, and when data is scaled it matches a model conditioned on the ground-truth context. On CartPole, the benefit of refinement grows with planning horizon: compounding error in the amortised model makes long-horizon planning worse than not planning at all, whereas planning through the refined model yields the largest reward gains we observe.
Chat is not available.
Successful Page Load