CARL: Cartridge Adaptation through Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards (RLVR) commonly adapts language models by updating model weights. This makes each specialization a new weight-space artifact and causes gradient and optimizer state to scale with model size. We ask whether useful reward-driven adaptation can instead be stored in a compact key–value (KV) prefix attached to an otherwise frozen model. CARL (Cartridge Adaptation through Reinforcement Learning) treats the cartridge itself as the policy parameters: rollouts are generated by a frozen base model conditioned on trainable per-layer KV states, task reward is computed from the output, and policy-gradient updates flow only to the cartridge. On Countdown with Qwen3-1.7B, a 128-slot cartridge containing 7.34M parameters raises held-out pass@1 from 11.0% to 46.5%. Under the same data split, prompts, rollout budget, and evaluation protocol, full-model GRPO reaches 53.5%. Thus the cartridge captures 83.5% of the observed full-model improvement while making only 0.43% as many parameters mutable and storing the learned adaptation in a 29.4 MB file rather than a 6.88 GB checkpoint. These single-seed results establish that online RL can write substantial task-adaptive behavior into persistent KV state; they do not establish superiority to weight-space adapters, and leave open matched LoRA, untrained-prefix, capacity, and multi-task controls.