Boundary-Aligned Caution Helps, but a Loss-CUSUM Gate Does Not: A Matched Audit of a Learning-Rate Gate for Continual LLM Post-Training
Zhaohui Wang
Abstract
Continual post-training of an LLM over shifting domains raises a control question a fixed schedule does not answer: at each step, keep updating at full rate, or slow down because the distribution just changed? A natural answer is to detect the change and lower the learning rate. We audit this on continual LoRA post-training of Qwen2.5-1.5B across five abruptly-switched domains with a CUSUM-on-loss gate ported from a prior reinforcement-learning controller. To separate the value of caution from the value of timing it, we run controls that fix the caution budget and rate multiplier and vary only where the slow-down is placed (5 seeds, 3 orders). (i) Placement carries value: an identical caution budget spent right after the ground-truth change points reduces forgetting relative to scattered random placement ($d_z=-1.0$, $p=0.002$)—the smallest $p$ in an explicitly enumerated 18-test family and the only one below the Bonferroni threshold (multiplicity-aware exploratory evidence, not confirmation). (ii) The gate aims but misses: its caution mask is boundary-biased (1.8× enrichment vs. 2.5× for the boundary-aligned control and 1.0× for random), yet at an exactly budget-matched comparison shows no detected advantage over random placement. (iii) And it does not beat a strong constant: it shows no detected difference from its integrated-rate-matched constant and loses on accuracy to a more cautious, action-matched constant—as does a tuned continuous schedule (FINCH). A fair evaluation thus needs budget/placement-matched controls and a strong constant, not a full-rate comparison.
Chat is not available.
Successful Page Load