LinMix: Data Mixing as Linear Domain Attribution
Abstract
Language model performance depends strongly on the mixture of domains in the pretraining data corpus. Popular data mixing methods select these domain weights by training many small proxy models on random mixes and then fitting a regression model to predict performance from the mix. We recognize this practice as a domain attribution problem, where the regression measures each domain's contribution to the target metric. Through this lens, the standard Dirichlet sampling of proxy data mixes is a poor regression design that requires tuning and cannot support a linear fit for hundreds of domains. We introduce LinMix, which instead samples mixes at the vertices of the capped simplex so that each domain is at its cap or absent. This hyperparameter-free design decorrelates domain effects, thereby enabling an interpretable linear model whose coefficients are the domain attributions. The optimal mix given by LinMix provably retains only the domains with positive attribution values and admits a closed form. Pretraining experiments with 1B-parameter OLMo2 and 4B Qwen3.5 language models on up to 288 domains show that LinMix outperforms RegMix and Olmix on reasoning and code benchmarks.