Decorative Backtracking: Verbalized Self-Correction Is Neither Necessary Nor Sufficient in a Reasoning Distill
Shrey Varma ⋅ Saihej Singh
Abstract
Interpretability infers what a structure *does* from what it responds to. Verbalized self-correction lets that inference be tested end to end: a reasoning model narrates reconsideration with "wait," "hmm," and "I made a mistake," and whether the reconsideration did anything is checkable against the answer. We test it causally in DeepSeek-R1-Distill-Llama-8B on MATH-500, using a force-answer probe that reads the model's *intended answer* at any mid-chain position and a harness in which every contrast is within-chain and same-truncation-point. **Necessity:** removing the verbalized correction leaves final-answer accuracy unchanged under SAE-feature knockout and under assumption-free text-level excision (equivalence-tested within $\pm 0.10$; $\mathrm{BF}_{01} = 3.5$–$4.7$ on the held-out test split). Difference-in-means direction ablation at three layers is the exception and *raises* accuracy, the opposite sign to what faithfulness predicts. **Sufficiency:** inserting synthetic corrections (three phrasings) at wrong-answer points, including into never-rescued failed chains where there is no recovery ceiling, rescues nothing ($9.5\%$ vs. $9.1\%$ baseline); clamping the internal carrier on is likewise flat, though that arm fails its manipulation check, so the sufficiency claim rests on the text result. The null is not an argmax artifact: the distribution over answers across shared seeds does not move either, in entropy, modal mass, or mass on the pre-correction answer. The markers are not only surface text: they coincide with a genuine residual-stream shift (paired $d_z \approx 0.70$–$0.76$ against position-matched sentence-start controls) whose magnitude and direction have chance-level AUROC ($0.44$–$0.49$) for predicting whether the answer actually changes, and which localizes revisions only weakly ($1.4\times$ their token share). Verbalized self-correction in this distill is a real internal event that is *decorative* with respect to the answer. A feature can be the best available detector of a behavior, and its span can mark a genuine state transition, while being causally inert for the behavior's function.
Chat is not available.
Successful Page Load