AutoDataBench: How Far Are LLM Agents from Autonomously Engineering Post-Training Data Pipelines?
Abstract
Post-training data pipelines are traditionally hand-designed by researchers who orchestrate generation, verification, and diversity balancing. While Large Language Models (LLMs) can already execute individual subtasks, the architectural composition of these pipelines remains a manual bottleneck, excluding the most substantial engineering step in post-training from automation. We introduce AutoDataBench, a benchmark evaluating whether LLM agents can autonomously design post-training data-synthesis pipelines. Each task gives the agent only a description of a target capability and asks it to produce a complete, executable pipeline whose generated corpus is then used to fine-tune a fixed student model and scored on the evaluation. Each task is delivered as a natural-language instruction and a sandboxed workspace given to the agent, paired with an evaluation protocol. We instantiate AutoDataBench on five capabilities: instruction following, competition math, deep search, repo-level software engineering, and terminal automation. For each task we pair the evaluation with a strong human-designed pipeline whose downstream score serves as a reference point for the agent. Experimental results reveal that frontier agents close most of the gap to the human reference on text-centric tasks (e.g., Opus-4.7 nearly matches humans in instruction following), yet a noticeable performance gap persists in tasks requiring executable-environment construction (SWE and Terminal). Furthermore, we observe a decoupling between downstream task-solving strength and pipeline-construction quality, where strong reasoning models like GPT-5.5 may still produce suboptimal pipelines. These findings suggest that autonomous environment synthesis remains the primary bottleneck for self-improving data loops, distinct from mere task-solving competence.