Exploring Lifelong Adaptation: In-Context Reinforcement Learning in Non-Stationary Environments
Abstract
Standard In-Context Reinforcement Learning (ICRL) typically assumes stationary dynamics within an interaction history, a restriction that limits its real-world applicability. We formalize Lifelong ICRL, a more realistic paradigm in which agents must continuously adapt to diverse non-stationary dynamics, ranging from abrupt random shifts to structured temporal drifts within a single lifetime. To probe the adaptation limits of this setting, we conduct a large-scale empirical study across non-stationary environments that span discrete symbolic reasoning and continuous physics-based control. We systematically investigate three core dimensions governing adaptation: (i) Training Regimes, comparing the generalization boundaries of domain randomization, stationary training, non-stationary training, and their staged combinatorial strategies; (ii) Model Architectures, evaluating modern sequence models, including Linear Attention and hybrid designs; and (iii) Impact of Scale, analyzing the relative contributions of interaction scale and model capacity. Our results reveal how these factors jointly shape robust in-context adaptation under unknown dynamics, yielding actionable insights for the design of generalist agents.