An Analytical Model of Compute-limited Multistage Training Pipelines
Abstract
Modern foundation models are trained through multistage pipelines, typically beginning with unsupervised pretraining, followed by supervised fine-tuning (SFT) on step-by-step solutions, and reinforcement learning (RL) with bulk feedback. Empirically, the design of these stages, particularly the compute allocated to each, strongly affects performance. Yet our theoretical understanding of multistage training remains limited. Here, we introduce a simple, analytically tractable model of multistage training. The model reproduces several qualitative phenomena: SFT and RL can be complementary, pretraining quality imposes a performance ceiling with power-law scaling, optimal compute allocation varies with data quality and pretraining level, and RL can selectively amplify task-relevant pretrained skills. Our model distills the complex behaviors into key ingredients, allowing clearer understanding of how and where they arise. Overall, our model provides a theoretical starting point for explaining the phenomenology of multistage training pipelines in the compute-limited regime.