Adoption-Gated Continual Learning for On-Device Tool-Calling SLMs
Jakub Mroz ⋅ Henry Ndubuaku ⋅ Karen Mosoyan ⋅ Parkirat Sandhu ⋅ Roman Shemet
Abstract
A small language model (SLM) that calls tools on a device receives one form of feedback: the user accepts or rejects each action. That signal is enough to build training data, but it does not say whether a model trained on that data should replace the one being served. We propose adoption-gated continual learning (AGCL), a simple loop that folds the adoption decision into learning. It keeps accepted and rejected behaviour in bounded sets and fine-tunes a candidate model on the accepted data. It scores both the candidate and the incumbent on two terms: how much accepted behaviour a model retains, and how much rejected behaviour it still repeats. A single exploration rate $\alpha$ weights the two terms, and the higher-scoring model is served. We run AGCL with a 270M-parameter model and full fine-tuning on three simulated domains, at $\alpha \in \{0, 0.5, 1\}$ with three seeds each. In all 27 preregistered runs, the served model ended at or above its starting accuracy on a held-out verification set. Moderate exploration improved accuracy where the conservative setting blocked adoption. Full exploration never beat moderate and discarded more accepted behaviour. A nearly saturated domain did not move at any $\alpha$.
Chat is not available.
Successful Page Load