Dynamic Sampling is Adaptive Data Collection: A Closed-Form Selection Bias that Curricularizes Safety for Free
Jonathon Romero ⋅ Grant C Forbes ⋅ Zixiang Tang ⋅ Adrian Chan ⋅ Yijin Wang ⋅ Sriman V Donthireddi
Abstract
DAPO's Dynamic Sampling is a prebuilt curriculum that inherently adapts during the training process. It discards any rollout group whose rewards are unanimous and resamples a replacement which is an adaptive data collection rule, and like any accept/reject rule it biases what the learner actually sees. We characterize that bias in closed form. Because contestedness falls as a domain is mastered, the filter reallocates gradient capacity away from mastered domains automatically: an inherent curriculum, with no reweighting term. The non-stationarity driving it is endogenous, with prompt pool, reward and mixture all held fixed, what reaches the optimizer still shifts, because mastery alone changes which samples survive selection. We validate the formula out of sample, predicting one run's realized batch composition from a second run's logs alone. Two controlled DAPO/Dr.\,GRPO fine-tunes of Gemma~4~E2B on mixed safety and coding prompts, identical but for the filter, extended to the same optimizer step count, show that the filter does not change whether safety is learned but what its cost is: both arms converge to the same trained-domain outcome, at $3.8\times$ the rollout compute, and, under a finite pool, a proportionally growing data cost. Unfiltered, capacity is not reallocated but attenuated in place, and continued pressure on a saturated domain produces over-refusal drift. Five external benchmarks confirm a real capability cost, IFEval falls $13.55$ points unfiltered against $2.66$ filtered at matched epochs, while GPQA-Diamond falls similarly in both, so the protection is specific to the axis the mechanism targets, not general.
Chat is not available.
Successful Page Load