RL Environments are an Infrastructure Problem: Lessons from Operating Agentic Rollouts at Scale
Abstract
Scaling reinforcement learning (RL) for agentic language models is a systems problem: every optimizer step needs thousands of concurrent, stateful, sandboxed environment executions to complete reliably. We report lessons from operating the rollout infrastructure of a production post-training stack as design recommendations for an OS layer for agentic AI. We read the environment interface as a question of who owns the rollout loop, and argue for capturing inference at a proxy gateway rather than standardizing an environment class hierarchy. The gateway contract is what makes unmodified agent applications (a.k.a. black-box applications) trainable and multi-agent structure visible to the trainer. Beneath the interface, sandbox provisioning is dominated by image economics and lifecycle, not isolation technology. Environments must stay dependency-light, with heavy tooling moved behind service boundaries. Infrastructure failures must be caught and labeled rather than scored, because a crash mistaken for a failed episode quietly spoils the training signal.