When Reasoning Meets Its Laws
Abstract
Despite the superior performance of Large Reasoning Models (LRMs), their reasoning behaviors are often counterintuitive, e.g., excessive reasoning on simple questions or insufficient reasoning on complex ones, leading to suboptimal reasoning capabilities. This paper presents the Laws of Reasoning (LoRe), a unified framework that formalizes intrinsic reasoning patterns in ideal LRMs. We first propose compute law with the hypothesis that the reasoning compute should scale linearly with question complexity. Beyond compute, we extend LoRe with a supplementary accuracy law, positing exponential decay of accuracy with increasing complexity. Since the complexity is difficult to quantify in practice, we approximate these hypotheses via two tractable properties, monotonicity and compositionality. We therefore introduce LoRe-Bench, a benchmark that systematically measures these properties for large reasoning models. Evaluation shows that most reasoning models exhibit reasonable monotonicity but lack compositionality. In response, we develop an effective finetuning approach that enforces compute-law compositionality. Extensive empirical studies demonstrate that better compliance with compute laws yields consistently improved reasoning performance on multiple benchmarks, and uncovers synergistic gains across properties and laws.