ADIS-Law: Unified Scaling Laws for Annealing-Phase Domain Injection in Large Language Models
Lyuxin Xue ⋅ Hui Cai ⋅ Xiaoyun Feng ⋅ Xuanwei Hu ⋅ Xin Zhang
Abstract
Continual pre-training (CPT) enables injecting domain-specific knowledge into large language models during the learning rate annealing phase. However, this compute-efficient paradigm forces practitioners to navigate a challenging optimization space: simultaneously choosing the injection fraction $\lambda$, the data mixture ratio $r$, and the re-warmup peak learning rate $\eta$. Because existing scaling laws have yet to comprehensively integrate and couple these interacting variables, they cannot reliably forecast the resulting trade-offs and final convergence outcomes across unseen injection configurations. To bridge this gap, we introduce the Annealing-phase Domain Injection Scaling Law (ADIS-Law), a parametric framework that explicitly models late-stage distribution shifts as physical perturbations applied to the base pre-training trajectory. ADIS-Law formalizes two opposing macroscopic behaviors: general-domain performance follows a linear forgetting-recovery process, whereas target-domain adaptation adheres to a power-law adaptation trend. Across data-budget and model-scale extrapolation, ADIS-Law predicts target-domain trajectories with an average Spearman correlation of $0.947$ and forecasts tail losses with relative error below $5\\%$. We further show that ADIS-Law can predict optimal injection strategies at extrapolated scales: it ranks target-scale candidate policies with a mean Spearman correlation of $0.936$. By replacing blind grid search over $(\lambda, r, \eta)$ with a single fitted law, ADIS-Law offers a data-driven approach for strategy selection in annealing-phase domain injection.
Chat is not available.
Successful Page Load