Anchoring Reasoning Distillation via Syntactic Constraints
Zehua Cheng ⋅ Wei Dai ⋅ Jiahao Sun
Abstract
Distilling "System 2" reasoning into compact student models is bottlenecked by the scarcity of fully-correct teacher traces: as task difficulty grows, rejection sampling discards an increasingly large fraction of teacher generations. "Wrong" (W) traces are abundant, but training on them indiscriminately causes *policy poisoning* the student internalises the teacher's hallucinated arithmetic alongside any useful reasoning structure. We propose **Atomic Reasoning Units (ARU)**, a framework that resolves this trade-off via strict syntactic constraints. ARU enforces a Backus--Naur-Form (BNF) grammar that decomposes each reasoning step into a Premise, an Operation, and a Result, algorithmically disentangling reasoning planning from arithmetic execution. *Counterfactual Logic Verification* (CLV) symbolically re-executes W-traces to recover the $68\%$ that have valid plans but faulty arithmetic, and *Syntactic Loss Masking* (SLM) trains the student on this verified structure while suppressing the gradient on hallucinated result tokens. Across the $3 \times 8 \times 6$ grid we evaluate ($\{0.6, 1.7, 8\}$\,B Qwen3 students $\times$ eight benchmarks $\times$ six baselines), ARU is the strongest method in every cell, improving GSM8K by $+11.1$ and MATH-500 by $+9.8$ over the closest baseline at $1.7$\,B; on the held-out AIME competition set ARU's Maj@$8$ is $14.8$ versus $9.5$ for PoT (the closest baseline). We further document a small-scale capability gap in which ARU at $0.5$\,B matches Gold-SFT at $\geq 1.5$\,B on GSM8K (paired-bootstrap $p<0.05$ after Bonferroni correction), with a pre-registered ProofWriter control showing the gap collapses on pure-logic reasoning---consistent with arithmetic offloading as the dominant mechanism, though we do not claim the control uniquely isolates it. The code is available at https://anonymous.4open.science/r/ARU-286D.
Chat is not available.
Successful Page Load