Which Environments Should LLMs Learn From? Predicting RL Utility Across Four Scales
Varun Madan ⋅ Brando Miranda ⋅ Zhanke Zhou ⋅ Jiarui Yan ⋅ Srivatsava Daruru ⋅ Elyas Obbad ⋅ Sanmi Koyejo
Abstract
Reinforcement learning with verifiable rewards can improve language-model reasoning. Yet equal compute can help, do little, or degrade a policy when training environments differ. We ask whether measurements from the starting policy can predict an environment's held-out utility before RL. We run 32 controlled Qwen3 trajectories from 1.7B to 14B parameters in a 2$\times$2 design that varies group-signal availability and the admission-cost profile. We evaluate 196 policies on a fixed, disjoint 1,000-prompt OpenR1 evaluation panel. Environment choice changes held-out pass@1 by up to 11.4 percentage points at matched generated-action-token budget. Reward variation alone does not explain this spread. We measure admitted advantage mass per generated token, $\rho$, and its effective coverage across prompts, $H$. A compact $\rho$-$H$ utility model is fit through 8B and frozen before inspecting 14B outcomes. At 14B, it achieves 2.29-point curve RMSE, 77.8\% pairwise rank accuracy, 2/3 correct best-environment decisions, and 0.29-point mean allocation regret. A four-scale leave-one-scale-out audit yields 2.09-point RMSE and 0.48-point regret. The $\rho$-$H$ model improves curve prediction over baselines based on cost and admission or on budget and scale. Together, these results provide evidence that pre-RL measurements help predict how environment choice changes held-out improvement across model scale and compute, providing a basis for more informed post-training compute allocation.
Chat is not available.
Successful Page Load