ISOPro: Controlling Verifiable Self-Training through Replay Composition and Task Coverage
Abstract
Post-training from model-generated data creates a feedback loop: the current policy determines which training examples the next policy receives. When verified successes are sparse, this loop can amplify sampling bias, discard useful capabilities, or spend additional inference without improving the training signal. We introduce ISOPro, a replay-controlled self-training method that samples candidate solutions, retains verified reasoning traces, and trains parameter-efficient adapters from replay buffers with explicit coverage policies. ISOPro separates three decisions that conventional self-training often conflates: which traces qualify as training data, how replay allocates capacity across task strata, and whether additional compute increases repeated sampling or task coverage. Across five seeds on a fixed 259-task scheduling evaluation, tier-balanced ISOPro improves greedy accuracy by 26.7 percentage points over the base model (95% hierarchical bootstrap CI: 20.5 to 33.0). Under matched selection budgets, verified selection outperforms random selection by 28.9 points (95% CI: 23.6 to 35.0), isolating trace correctness as a causal source of improvement. Increasing rollouts per task from one to eight produces gains between 21.1 and 25.0 points, whereas increasing coverage from one to two tasks per tier raises the gain from 7.7 to 26.6 points. A controlled five-seed comparison finds that ISOPro outperforms the tested GRPO baseline by 24.9 points (95% CI: 18.9 to 31.4), and a separate 7B replication improves by 18.0 points over its base model. These results identify replay composition and task coverage as central controls for post-training from sparse verifiable feedback.