How New Strategies Emerge in RL Post-Training: A Controlled Study
Abstract
Does RL post-training build new reasoning procedures, or merely reweight behaviors already latent in the base model? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite can be audited. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained with only a binary final-answer reward. RL solves held-out problems that remain rarely solved by the pretrained model even under much larger sampling budgets, while rejection fine-tuning improves early but plateaus. Trace analysis reveals a phased mechanism: RL first strengthens primitive reductions, then enters a chunking phase in which it forms valid compressed procedures---macro contractions that collapse sequential reductions and parallel contractions that combine independent ones. These procedures are not isolated samples; they are reused and consolidated into a stable repertoire. Comparing RL with rejection fine-tuning shows that the key difference is not exploration volume but selectivity: RFT produces many shortcut-like rewrites, much of them invalid, whereas RL concentrates exploration into valid reusable structure. Pretraining ablations show that strategy emergence is gated not by primitive exposure alone, but by whether pretraining organizes primitive competence into reduction procedures that RL can later compress. The base model provides weak procedural ingredients; RL builds them into reliable higher-level strategies.