Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery
Abstract
Large Reasoning Models (LRMs) generate long reasoning traces before producing an answer, so deploying one under a device memory budget calls for aggressive weight quantization. However, at 2 bits compression breaks models in qualitatively different ways that require different solutions. We study full reasoning traces under 2-bit quantization and find two distinct failure modes: path-finding failure, where the model loops over unproductive steps without ever reaching a valid answer, and commitment failure, where it reaches a correct answer mid-trace but keeps going instead of stopping. Both inflate latency and exhaust the generation budget, so the per-token gain does not survive end to end. We propose one targeted intervention for each: FP16 planning, a high-precision outline that keeps the 2-bit model on track, and loop rescue, which detects repetition and either commits to the current answer or reruns the example at full precision. On Qwen3-8B, loop rescue recovers MATH-500 accuracy from 17.2% to 74.2% and reduces average trace length by 92%. On Qwen3-32B, combining both interventions closes 22.2 accuracy points over direct 2-bit execution. Our code is available at: https://github.com/brain-lab-research/quantized-reasoning/.