Induction and Adaptation in Natural Language
Abstract
A general AI agent cannot anticipate every environment it will encounter and must therefore continually adapt its understanding of the world based on observations made at deployment. We investigate the benefits of representing and compressing such beliefs in natural language, which, paired with a pre-trained language model, forms a highly flexible world model that can map states and actions to predicted next states. Unlike the scalar signals that drive traditional neural world models, natural language is interpretable, editable, and portable. We formalise the inference of such a world model as semi-amortised inference over a text-valued latent variable. To search within the natural-language latent space, we use a language model as a proposal that generates hypotheses from a strong prior, grades them against observations, and refines them in response to prediction errors. We evaluate our method on ARC-AGI-3 and AutumnBench, two benchmarks that require agents to explore complex environments and infer their dynamics to solve novel tasks. Our method substantially outperforms in-context learning baselines on next-state prediction, and we find that the induced natural-language world models serve as effective representations that drive strong downstream task performance.