When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment
Abstract
The reasoning trajectory of the Large Language Model (LLM) is often regarded as the verbalized description of the internal thinking. However, the unfaithfulness of the reasoning process introduces the risk of shortcut reasoning, where the model fails to reason step by step but instead relies on discovered shortcuts to reach the final answer, while post-rationalizing this decision through a seemingly coherent verbalized reasoning process. This shortcut reasoning is difficult to detect, as existing monitors and verifiers mainly inspect textual reasoning traces or final outcomes, failing to capture how the model’s answer belief forms during generation. To figure out the intrinsic pattern in shortcut reasoning, we propose ConfLens, a framework that tracks how a model's confidence in its final answer evolves throughout the reasoning process. Across three shortcut reasoning settings, we find that shortcut samples often exhibit premature confidence, characterized by high confidence in the final answer at early reasoning stages. However, reliably detecting this pattern remains challenging, as existing confidence estimation methods are limited in generalizability, reliability, and efficiency. To address this, we introduce the Distributional Answer Commitment Score (DACS), a distributional confidence estimation method that instantiates ConfLens for effective shortcut reasoning detection. DACS estimates the entropy of the model's probability distribution over answer commitment during reasoning, thereby capturing how concentrated the model's answer belief is at each reasoning step without access to the ground-truth or task-specific verifier. We further convert the detection results of ConfLens into interpretable signals to mitigate reward models’ preference for shortcut reasoning responses. Experiments on math and code reasoning tasks show that ConfLens instantiated with DACS improves shortcut reasoning detection by over 4.3% in F1 compared with strong baselines, while reducing the gap between faithfulness and correctness in reward model preferences.