HALO: Homotopy-Augmented Layer Optimization for Stable LLM Supervised Post-training
Abstract
Supervised post-training, including supervised fine-tuning (SFT) and knowledge distillation (KD), is essential for adapting large language models (LLMs) to specialized domains. However, traditional supervised post-training methods often suffer from catastrophic forgetting, training instability, and prohibitive computational costs. While layer-wise strategies have emerged as efficient alternatives, they introduce a representation bottleneck, where constrained updates in early stages limit feature diversity and impair final generalization. To resolve these issues, we develop HALO (Homotopy-Augmented Layer Optimization), a homotopy-driven layer-wise strategy for LLM post-training. By introducing a continuous homotopy parameter and a sequence of monotonically increasing auxiliary functions, we formulate the post-training process as a continuous deformation of the optimization problem that progressively incorporates layers into the trainable set. Specifically, HALO unfreezes the model from the output back to the input in a structured sequence. We theoretically prove that this approach establishes a differentiable solution path, ensuring a stable transition from a restricted optimization state to a fully converged LLM. Extensive experiments across diverse model families and sizes demonstrate that HALO consistently enhances performance and generalization. Notably, HALO exhibits strong robustness in challenging scenarios, especially in low-resource settings, while significantly reducing convergence time and iteration counts. Our results position HALO as an effective and accessible framework for high-performing LLM specialization.