Correctable Fork Tokens: Verifier-Anchored Selective Credit Assignment for Tool-Integrated RLVR
Shuqi Yin
Abstract
In tool-integrated reinforcement learning with verifiable rewards (RLVR), sequence-level verifier signals are issued for trajectories whose outcomes are determined by a sparse subset of tokens—tool names, argument keys, and argument values. Standard RLVR assigns the same advantage to every token in a rollout, diluting credit at precisely the positions that matter. We propose **Correctable Fork Tokens (CFT)**, which decomposes selective credit assignment into two subproblems: *where* to update, and *how much* to update at those positions. To this end, CFT introduces a training-only answer-conditioned branch sharing parameters with the student policy. The key insight is that conditioning on the ground-truth answer concentrates the model's distribution at decision-fork positions, providing hindsight localization without additional parameter capacity or imitation targets. Critically, this benefit requires accurate answer information: a larger-capacity teacher without the answer fails to replicate these gains, confirming that privileged answer information—not teacher capacity—is the active ingredient. CFT uses this entropy reduction ($IG_t$) to identify correctable fork positions, then rescales token-level credit at those positions while preserving the verifier advantage as the sole signal for update direction. On BFCL v3/v4, CFT improves multi-turn accuracy by **4.2/6.5** pp and overall accuracy by **3.1/2.2** pp over GRPO; it further surpasses On-policy distillation—which employs a larger Expert Model as teacher—by **3.6/6.7** pp on multi-turn, confirming that answer-conditioned fork localization rather than teacher capacity drives the improvement. On Tau3, the average score improves by **8.3** pp. We additionally present a three-level structured verifier for tool-call settings; ablations confirm that finer-grained verifier feedback and selective token credit are complementary. Our reproducible implementation is available at [https://anonymous.4open.science/r/CFT](https://anonymous.4open.science/r/CFT).
Chat is not available.
Successful Page Load