Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary
Dane Malenfant
Abstract
Agents that learn alongside one another face a moving target: each agent's environment shifts as its partners update. We study how long success-conditioned reusable structure survives this drift. Our starting point is the standard focal-agent induced-MDP representation: marginalizing peers whose policies are fixed within an episode preserves every focal trajectory law and expected return, and an evolving peer produces the continual sequence of MDPs faced by the focal agent. An \emph{invariant core} represents reusable structure through maximal abstract patterns appearing in a high fraction of successful focal trajectories. Our main result is a worst-case-tight conditioning theorem: trajectory-law drift $\varepsilon$ can reduce a candidate's success-conditioned coverage by at most $\frac{\varepsilon}{p_0}$, where $p_0$ is its reference success mass, and the coefficient is sharp. Peer-policy movement supplies $\varepsilon$; positive coverage margin then yields a certified $\Omega(\frac{1}{\eta})$ survival horizon and, under an explicit effective-conflict condition realized by exact policy gradient in an analytic class, a matching $\Theta(\frac{1}{\eta})$ first-exit law. With calibrated success mass and executability, the same certificate yields policy-value, library-selection, and transfer-regret guarantees. An exactly solvable corridor confirms the structural predictions, including the inverse-rate lifetime ($R^2>0.9999$). Two registered 64-stream studies in continual control and cue-MNIST show that core erosion predicts impending failure and enables near-oracle intervention; an exploratory reanalysis of eight learned-partner Level-Based Foraging development pairings suggests the same erosion--failure link under peer learning.
Chat is not available.
Successful Page Load