A Computational Perspective to Data Ablation Experiments
Abstract
\emph{Scaling} has been the central driver of progress in modern AI, yet data curation remains an exception: it still relies on manual heuristics rather than systematic, scalable methods. In practice, data recipes (e.g., filter thresholds, deduplication strength, domain mixing ratios) are determined through many small-scale ablation training runs. This work demonstrates that the performance of the best-discovered recipe improves \emph{predictably}--as a \emph{power law}--in the number of such ablation runs. Theoretically, the power-law form arises naturally from regret bounds in zero-order optimization theory. Empirically, we validate it across diverse training stages, including pretraining, supervised fine-tuning, preference optimization, and reinforcement learning. Building on this finding, we derive a principled approach to \emph{data quality--quantity tradeoff} under a fixed compute budget, yielding a Pareto frontier that dominates standard heuristic pipelines (e.g., DCLM). Together, this work provides a general framework for scaling data-curation compute, establishing it as a new axis alongside the well-studied pretraining and test-time scaling.