LLMs Keep Thinking When Told Not To
Abstract
Large Language Models (LLMs) increasingly ship with “thinking modes”, yet their counterpart, the “no-thinking” has received far less exploration. In this work, we revisit language model “no-thinking” behavior along the following efforts: (a) Explore the definition of no-thinking. Rather than relying on proxies like thinking-mode control, we split each response into pre-answer text and a final answer, scoring both answer-only compliance and the question–pre-answer relevance. (b) Broad experimental coverage. We evaluate multiple no-think controls (thinking-mode toggles, prompt instructions, etc.) across boolean, multiple-choice, and open-ended tasks, on open-source models with scaled variants, and proprietary models. With extensive experiments, we found that (i) Explicit no-think controls do not reliably eliminate visible reasoning-like content. Instead, models exhibit (ii) “thinking inertia”: residual content is systematically organized by the task's answer space, with boolean verification compressing most easily, multiple-choice tasks forming an intermediate regime, and open-ended tasks remaining the most resistant. (iii) We further identify what we call a no-thinking and performance trade-off in open-ended: stronger answer-only constraints on complex tasks either fail to suppress the intermediate payload or succeed by removing computation needed for accuracy. (iv) Finally, we show that compressibility depends strongly on the structural support provided by the question itself. Thus, no-think controls are not direct guarantees of no-thinking; they must be evaluated jointly through answer-only success, visible payload, and task accuracy. We hope our work moves LLMs closer to a human-like ability to stop thinking on demand.