How Data Scales in Agentic Reinforcement Learning: Laws and Synthesis Strategies
Bowei He ⋅ Yankai Chen ⋅ Xiaokun Zhang ⋅ Changjiang Han ⋅ Ye Yuan ⋅ Chong Li ⋅ Steve Liu
Abstract
Pretraining scaling laws treat training data as a one-dimensional quantity: a token count. Agentic reinforcement learning (RL) inherits this scalar abstraction, but modern synthesis pipelines now generate *tasks*, *environments*, and *trajectories* independently and at very different costs, turning data into a structured design space. This raises a question pretraining never had to answer: given a synthesis budget, *which axis* should be scaled? We address this question end-to-end with three tightly coupled contributions. First, we decompose agentic data into four axes (task, environment, trajectory, reward) and empirically fit per-axis scaling laws on Qwen3 models (4B/8B/14B) across math, code, and web/UI domains; we observe a robust ordering of per-axis scaling exponents that holds across model sizes, RL algorithms (GRPO and PPO), and domains. Second, we use the fitted laws as a common ruler to price synthesis methods, capturing each method's efficiency as a per-axis discount factor and identifying critical synthetic-to-real ratios at which collapse begins; these ratios differ by an order of magnitude across axes. Third, we cast budget allocation as a constrained optimization under our laws and discounts, deriving a Chinchilla-style compute-optimal recipe and validating it at held-out scale. The empirical results demonstrate that following the recipe yields up to $1.7\times$ compute efficiency over the dominant ``scale-the-trajectories'' practice.
Chat is not available.
Successful Page Load