Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
Abstract
Large Reasoning Language Models (LRLMs) leverage long Chain-of-Thought (CoT) reasoning to solve complex tasks, yet often engage in overthinking. Specifically, they continue to generate reasoning steps that provide meaningless contribution to or even degrade the final answer, even after sufficient reasoning has been produced. Early-exit strategies are proposed to dynamically terminate meaningless or even harmful reasoning, aiming to improve both efficiency and accuracy. However, existing methods either suffer from over-truncation that degrades accuracy below vanilla CoT, or yield only spurious efficiency gains, which reduce token counts but increase wall-clock inference time. We observe that when an LRLM's reasoning path deviates into meaningless or even harmful wandering, this deviation is accompanied by an anomalous surge of high-entropy transition tokens. Building on this insight, we propose RPDI-EE, a training-free early-exit method that monitors the Reasoning Path Deviation Index (RPDI), defined as the ratio of local to global average entropy, which serves as a proxy for reasoning path deviation. By tracking this index, RPDI-EE detects and terminates likely meaningless or even harmful reasoning while preserving productive inference. Experiments across multiple benchmarks and four LRLMs of varying types and scales show that RPDI-EE achieves the largest average accuracy improvement over vanilla CoT among all tested early-exit methods, while mitigating the spurious efficiency gains of existing alternatives.