Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models
Abstract
Capability and efficiency are two key dimensions of reasoning in large language models (LLMs). Capability refers to the ability to solve a problem correctly. We operationalize efficiency as output length among correctly solved instances, with shorter outputs indicating greater token efficiency. When LLMs use chain-of-thought reasoning to solve problems of controlled hardness, how accuracy and output length vary jointly with hardness and model size remains poorly understood. Here, we use hierarchical Bayesian models to evaluate LLMs from the DeepSeek-R1-Distill model family across four arithmetic and algorithmic reasoning tasks. At a fixed model size, the probability of correctly solving an instance decays approximately exponentially toward an asymptote with instance size, our proxy for problem hardness. The inferred decay scale grows sublinearly with model size, indicating that capability increases but with diminishing gains. Among correctly solved instances, output length grows as a power law with instance size. We find little evidence that the parameters of this power law vary systematically with model size. Our findings reveal potential limitations of naive model-size scaling: capability improves with diminishing returns, while token efficiency shows little evidence of systematic improvement.