Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Flow
Abstract
Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, time spent improving the upstream objective reduces the time available for downstream adaptation. Despite its practical importance, this trade-off is not yet well understood theoretically, even in simple models. In this paper, we cast compute allocation as a time-split problem under a two-stage pretrain–fine-tune procedure with fixed total optimisation time, using regularised least squares trained by gradient flow as a tractable setting. We characterise the optimal switching time under data-dependent evaluation geometries induced by the fine-tuning problem. Our results show that the allocation depends on how pretraining directions affect fine-tuning predictions and how fine-tuning shifts are seen through downstream data geometry. In particular, the relevant quantities are determined by prediction-relevant spectral components of the pretraining and fine-tuning empirical covariances. Technically, the analysis relies on a basis-invariant, eigenspace-level spectral decomposition, together with perturbative control of the non-commuting pretraining and fine-tuning dynamics.