Contrastive Thinking Decoding: Steering Answer Generation in Reasoning Models
Taehyeon Kim ⋅ Youngsoo Jang ⋅ Hyunsoo Lee ⋅ Yu Jin Kim ⋅ Moontae Lee
Abstract
Large reasoning models separate inference into an explicit thinking phase followed by a final answer phase, yet how the answer phase uses the reasoning trace remains unclear. We present a systematic study of this trace-to-answer transition across multiple LRMs and find that (i) final answers diverge from correct reasoning traces at rates between 6\% and 48\%, even when traces terminate cleanly at full budget, (ii) naively copying the trace answer underperforms standard decoding in some settings, revealing that the answer phase has independent corrective capacity, and (iii) the same phenomenon extends beyond structured tasks to settings where reasoning traces and final answers can diverge under social-pressure framing (sycophancy). These findings motivate Contrastive Thinking Decoding (CTD), a training-free, single-model decoding method that contrasts answer-phase logits under the primary reasoning trace against those from a deliberately degraded noisy trace. CTD selectively amplifies trace-consistent tokens while preserving the answer phase's corrective capacity, concentrating CTD-internal joint mass on the both-correct cell (T$\checkmark$A$\checkmark$) without inducing trace-correct-to-answer-wrong drift relative to the base distribution. Across math reasoning, science, code, social-pressure resistance (sycophancy), and a broad knowledge benchmark (MMLU), \ctd reduces trace-answer disagreement and improves the accuracy--compute trade-off without parameter updates or auxiliary models at decoding time.
Chat is not available.
Successful Page Load