It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Pierre-Carl Langlais ⋅ Pieter Delobelle ⋅ Yannick Detrois ⋅ Pavel Chizhov ⋅ Carlos Rosas-Hinostroza ⋅ Neil S Smail ⋅ Benjamin Burtin ⋅ Hanna Shcharbakova ⋅ Ivan Yamshchikov ⋅ Anastasia Stasenko
Abstract
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines—for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, e.g., with reasoning traces to address cold start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called _synthetic data_ on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present Synth, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate Synth by training a series of models: a 56M tiny model (Monad), 0.3B--0.6B dense models (Baguettotron), and a 13B-total / 1B-active Mixture-of-Experts. At iso-compute, Synth outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because Synth is back-translated from grounded passages, Synth-trained models achieve high factual precision despite 10-140$\times$ fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our Synth dataset, are capable of producing competitive generalist models at a significantly lower cost, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our Synth dataset and the series of Baguettotron models under a permissive license, thus supporting open-source language model development.
Chat is not available.
Successful Page Load