Dream-RSI: Off-Policy Meta-Learning for Recursive Self-Improvement
Abstract
AI agents are increasingly applied to open-ended scientific discovery, where progress depends on effective exploration over large and complex search spaces. However, selecting and improving exploration strategies is itself a challenging meta-learning problem. As discovery scales, the space of possible exploration strategies quickly becomes too large for manual design, while evaluating such strategies requires costly long-horizon rollouts with delayed and expensive feedback. We introduce \textsc{Dream-RSI}, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By replaying trajectories generated by previous controllers, \textsc{Dream-RSI} obtains inexpensive off-policy feedback for evaluating and improving the current exploration controller without repeatedly rerunning costly discovery. The improved controller is then redeployed online to guide further discovery, generating new experience for the next round of improvement. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, \textsc{Dream-RSI} achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.