Base Items Overfit, New Items Underfit: Hidden Cost of Joint Training in Incremental Adaptation
Abstract
Pretrained sequential recommenders and language models are routinely extended with new entries by jointly fine-tuning old and new embeddings. We show this has a hidden failure mode: old-entry quality degrades while new entries still improve, forcing premature early-stopping. We propose population-specific low-rank subspaces: base and new entries are parameterized with separate shared projection matrices, decoupling their optimization while providing implicit regularization proportional to data availability. Three instantiations (Freeze-SV, Freeze1-SV, Dual-SV) prevent the failure while maintaining or improving new-entry quality. On sequential recommendation (two architectures, two large-scale datasets) our methods Pareto-dominate joint fine-tuning and continual-learning baselines (EWC, ADER) on the base-vs-new quality tradeoff. On LLM vocabulary expansion (three model sizes across two domains), at least one frozen variant wins on overall perplexity in 8 of 9 model-scale cells on both domains, and our low-rank parameterization matches or beats full-rank quality at 4× fewer new-entry parameters (r/d = 25%).