Benchmarking World Models for Continual Learning on Compositional Tasks
Abstract
A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what it has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences underpins an agent's ability to efficiently adapt to novel environments, as the dynamics of the physical world can often be described in reoccurring mechanisms. Their measure of adaptation, however, entangles two abilities: the speed and capacity to learn unseen tasks, and the reuse of knowledge already acquired, because incoming tasks always carry novel content alongside what recurs. In order to isolate knowledge reuse from prior experiences, we propose a compositional continual learning benchmark for world models in robot manipulation. Specifically, we design each task curriculum to end with a compositional task built by recombining all primitives that precede it in the sequence. We further factorise this composition along the axes of action and perception to better understand how different input modalities bottleneck knowledge reuse. We evaluate state-of-the-art world models under canonical continual learning methods, alongside a modular world model whose dynamics backbone contains explicitly reusable components. Results show that modularity balances reuse against forgetting better than conventional methods, but none solve the problem fully, leaving clear room for continual world models built to reuse without forgetting.