Concise Reasoning Through the Lens of Lagrangian Optimization
Abstract
Concise reasoning in large language models (LLMs) seeks to generate only essential steps needed to arrive at a final answer, thereby alleviating issues of overthinking. Most proposed approaches scalarize length and reward into a single objective, requiring coefficients or thresholds that must be re-tuned across domains and model scales. We address this brittleness by treating concise reasoning as a constrained problem that minimizes length subject to an accuracy floor, and deriving a tractable algorithm, Performance-Aware Length Update (PALU), which replaces each intractable update of Lagrangian optimization with a tractable surrogate while preserving its structural prescription. On DeepSeek-R1-Distill-Qwen-1.5B, PALU cuts generation length by 64\% and improves accuracy from 44.2\% to 51.6\% across six benchmarks, a level matched by GRPO only with 2.5× more tokens. The Lagrangian dynamics yields a two-phase compression: an initial phase where the residual signal still allows compression without performance cost, followed by a trade-off phase where performance gates further length reduction. The same hyperparameters transfer across math, logic and STEM, and across 1.5B to 14B models, suggesting that constrained optimization is a productive design scaffold for training reasoning LLMs.