DoubtLess: Training-Free Contrastive Decoding for Efficient Mathematical Reasoning
Abstract
Reasoning models earn their accuracy with long chains of thought, much of which re-verifies an answer that has already been reached. Existing ways to shorten the chain either train the model, with length penalties or fine-tuning on concise traces, which has to be repeated for every model, or apply a fixed rule at decoding time: stopping early, masking hesitation words, or asking for brevity in the prompt. We propose DoubtLess, a training-free method in which a much smaller model of the same family steers the large model's logits. The small model reads the same prefix as the large model twice, once as a brief thinker asked to go straight to the answer and once as a verbose thinker asked to double-check everything; the difference between the two logit vectors, scaled by a strength α, is added to the large model's logits inside the thinking block. The small model thus supplies only a direction over the shared vocabulary, obtained from nothing but the two instructions, while the large model keeps the candidate set, the stopping decision and the answer; the cost is 4% of the FLOPs per token. On Qwen3-32B, with a 0.6B model supplying the contrast, DoubtLess removes 37–52% of the thinking tokens on MATH-500, GPQA-Diamond, MMLU-Pro and BBH with accuracy within one point of plain decoding, and 27–34% on five competition-mathematics benchmarks under sampling with accuracy no lower than plain decoding on any of them; the strength gives a continuous trade-off between length and accuracy. As a control, we apply the same logit adjustments to randomly chosen tokens: the compression disappears, so the effect comes from the tokens the small model selects. Most of the removed tokens follow the first statement of the answer.