DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models
Abstract
Data-centric training—selecting, mixing, and reweighting data during LLM optimization—has produced a rapidly growing body of algorithms, yet every new method ships as a bespoke codebase with its own interface, training protocol, and evaluation setup, making head-to-head comparison and reproducible evaluation nearly impossible. We present DataFlex, the first unified benchmark and evaluation platform for data-centric training of LLMs. DataFlex contributes: (i) a standardized evaluation protocol that runs selection, mixture, and reweighting methods under identical seeds, data splits, hardware, and metrics; (ii) a reference implementation of ten representative algorithms (six data selection, three data mixture, one data reweighting) re-engineered on top of LLaMA-Factory to share model, dataloader, and distributed-training infrastructure (including DeepSpeed ZeRO-3 with full-rank gradient acquisition); and (iii) an empirical benchmark whose head-to-head results yield reproducible findings that no single existing codebase could produce: dynamic selection methods consistently outperform static full-data training on MMLU across Mistral-7B and Llama-3.2-3B, and the gain is larger on smaller models (up to 13.3 points); DoReMi and ODM yield complementary wins on high- vs. low-resource domains when pretraining Qwen2.5-1.5B on SlimPajama at 6B and 30B tokens; and DataFlex's implementation is consistently faster than the original codebases of LESS, TSDS, and MoE-SFT. By unifying fragmented methods under a single evaluation substrate, DataFlex provides the community with a reproducible, extensible basis for future data-centric research.