Adapt, Merge, Validate: Zero-Replay Continual Learning for Multi-Policy Safety Classifiers
Abstract
Production content-safety classifiers must be updated continually as inputs drift and policies evolve, yet after every update the model must perform well on every policy (in our study, we track recall at a fixed precision for individual policies). The standard remedies, e.g., retraining from scratch on the augmented data pool or fine-tuning with a replay buffer of old data, tie every update to a training job and to persisting sensitive user data, which carries compliance and privacy requirements. They also scale poorly: as pools grow, a fixed buffer covers little of the original distribution, and every adjustment of the data mix costs another full training run. We study a compute-efficient alternative for multi-headed safety classifiers based on a Gemma-3 encoder: the model is fine-tuned only on the new data, and its weights are interpolated with the previous model. Training uses only the new data, so no prior data is retained. Performance on existing datasets is maintained by merging with the previous checkpoint. The trade-off between new and existing datasets is tuned using a single coefficient chosen after training. We benchmark this method against replay-based fine-tuning across different buffer sizes. On the three shifts we studied, the outcome depends on the type of shift. On boundary shifts, the selected merge outperforms replay-based fine-tuning across all buffer sizes. On distributional shifts (i.e., novel inputs), replay-based fine-tuning performs better but exhibits an erratic frontier across buffer sizes; in contrast, pre-merge checkpoint averaging provides a consistently stable trade-off curve. The coefficient for the optimal merge differs across different shifts and must be empirically chosen. We also demonstrate that adapting to one distribution at a time and merging sequentially outperforms learning the two distributions jointly. Across sequential merges, we show that operating-point recall degrades at a rate governed by the strictness of per-policy precision targets. This yields a simple rule: merge while performance per policy stays above its threshold, and retrain when one falls below. Because each update reduces to one fine-tune, one coefficient, and one threshold determination, we implement the loop as a long-running, stateful agent running autonomously between two human sign-off gates, one before data generation and one before the final model is chosen.