PeerConf: Self-Calibrating Early Stopping for Efficient Parallel LLM Reasoning
Abstract
Sampling multiple reasoning traces in parallel significantly improves Large Language Model (LLM) accuracy on complex mathematical problem-solving, yet standard self-consistency runs every trace to completion regardless of quality or early resolution. While confidence-based filtering can mitigate this, current methods rely on a dedicated warm-up phase that wastes significant compute before any filtering begins. We introduce PeerConf (Peer Think with Confidence), a warm-up-free, self-calibrating confidence filter tailored for efficient test-time scaling in mathematical reasoning. PeerConf estimates percentile thresholds dynamically from live-finishing traces within the run itself, allowing filtering and consensus checks to begin immediately at the first finisher. Furthermore, PeerConf introduces reasoning-closure completion probes that test active traces; a trace that clears a confidence threshold and successfully closes its own mathematical reasoning commits early, casts its answer as a vote, and frees its compute seat. Evaluated on challenging competition mathematics benchmarks, including MATH500, AIME25, and HMMT25 using DeepSeek-R1-8B and GPT-OSS-20B, PeerConf matches or exceeds self-consistency accuracy while reducing token consumption by 53.1% to 79.8% relative to self-consistency at 32 traces, outperforming existing warm-up methods in token reduction. By removing the need for separate calibration phases, PeerConf makes extensive test-time scaling practical for resource-constrained mathematical reasoning applications. Code and benchmarks are available at https://github.com/mofe-stack/PeerConf.