PROWL-MARL: Active Failure-Driven Continual Learning for Multi-Agent World Models
Abstract
World models allow agents to learn from imagined experience, but their usefulness depends on how accurately they predict the parts of the environment that matter for decision making. Continually training a world model on new experience does not guarantee that its important mistakes will be corrected, because most new data comes from what the current task-solving policy already does. We introduce PROWL-MARL, an active continual-learning framework that adds a developer policy whose role is to find trajectories where the current world model predicts poorly. These failures are measured against observed environment continuations and prioritized for subsequent world-model updates, while leaving the underlying model architecture and learning objective unchanged. We evaluate PROWL-MARL on six procedurally varying StarCraft Multi-Agent Challenge v2 (SMACv2) scenarios, where model fidelity is measured through unit deaths, legal-action masks, rewards, and observations. Targeted repair consistently reduces error on discovered failures and also improves held-out reward and observation prediction across all six scenarios, while death and action-mask prediction show more task-dependent effects. These improvements also translate to stronger downstream policy learning across the benchmark, although the size of the gain varies across scenarios. Our results show that continual world-model learning can benefit from actively deciding which new experience should be used to improve the model.