What Does Language Model Overtraining Reveal About Deployment Objectives?
Justin Hartenstein ⋅ Andreas Haupt ⋅ Sanmi Koyejo
Abstract
Compute-optimal scaling laws divide a training budget between parameters and tokens, yet released language models often depart from pretraining-optimal choices: small models are trained with more tokens per parameter than their larger siblings. Among 32 comparable dense models from five developers, training tokens scale with parameters to the power $0.226$ ($\pm 0.130$ generation-clustered), far below proportionality. We model these allocations as creator choices that balancing desired use and training cost. For the Meta Llama 3 family of models, we find that models are overtrained by $114.8T$--$150.9T$ tokens for the 8B model and $33.0T$--$55.0T$ for the 70B model. We can reject a common data price and and identical distribution objectives. While these estimates remain illustrative as the released models are much smaller than current frontier models, both reveal a greater marginal value of compactness for smaller models. Our code is available at https://anonymous.4open.science/r/overtraining-objectives-9BA2/README.md.
Chat is not available.
Successful Page Load