Task Vector Descent: Learning from Non-IID Batches
Anton Baumann ⋅ Jonas Hübotter ⋅ Zeynep Akata ⋅ Andreas Krause
Abstract
Language-model adaptation data often comes from multiple domain-, user-, or task-specific distributions. Rather than being sampled i.i.d. from their overall mixture, successive minibatches may be temporally clustered by distribution. We study how the length of these same-distribution sequences affects the balance between learning from the targeted distribution and retaining performance on the non-targeted distributions. We find that longer sequences improve performance on the targeted distribution but degrade it on the non-targeted distributions, increasing the risk of catastrophic forgetting. We ask whether fully integrating the update produced by each sequence is the most effective way to learn from temporally clustered distributions. We compare full integration ($\lambda=1$) with partial integration, which applies a fraction $\lambda$ of the resulting parameter displacement (the task vector) to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, intermediate values of $\lambda$ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In a controlled continual-pretraining comparison, task-vector scaling outperforms reducing the learning rate by the same factor and fully integrating the resulting update, showing that its benefits are not reproduced by learning-rate scaling alone.
Chat is not available.
Successful Page Load