Greedy Local Learning for Language Model Pretraining: Gaps and Objective Design
Jihwan Moon ⋅ Sheir A. Zaheer ⋅ Jinmyoung Lee ⋅ Gunhee Kim ⋅ Chan Youn Park
Abstract
Greedy block-wise local learning splits a network into gradient-isolated blocks trained by local auxiliary losses, deleting the backward pass between blocks: inter-stage communication becomes forward-only and every block can step its optimizer independently, properties directly relevant to decentralized model-parallel training. Local learning is competitive with end-to-end backpropagation on image classification, but it has not been studied for autoregressive language model (LM) pretraining under the standard next-token objective with likelihood evaluation. We present a token-budget-matched empirical study at 125M and 400M parameters with $K\in\{1,2,4\}$ blocks at Chinchilla-optimal budgets, factorizing the auxiliary design into network architecture and training objective. We observe: (i) the gap to end-to-end training more than doubles from $K{=}2$ to $K{=}4$, but at $K{=}4$ shrinks from 125M to 400M; (ii) replacing an MLP auxiliary with an attention-bearing one is a strong network-side intervention, recovering 22--41\% of the gap; (iii) a multi-token-prediction (MTP) auxiliary objective helps at the first block boundary, whereas adding it at deeper boundaries hurts, and restricting it to the first block yields the best $K{=}4$ configuration ($+0.062$ vs.\ $+0.075$ nats at 400M); and (iv) deployment-style per-block execution reduces activation memory by up to $2.2\times$. We frame these results as an empirically grounded method direction rather than a finalized method: local objectives should apply future-predictive pressure selectively across boundaries while resisting shortcuts that bypass predictive content.
Chat is not available.
Successful Page Load