PeerConf: Self-Calibrating Early Stopping for Efficient Parallel LLM Reasoning
Abstract
Sampling many reasoning traces in parallel and voting is known to improve accuracy in Large Language Models (LLMs), but every trace runs to completion regardless of quality. How much inference-time computation a parallel run requires, and when it can be halted, is determined by signals that emerge only as traces finish, rather than by quantities fixed before the run begins. Confidence-based filtering terminates weak traces early, but current methods, like Deep Think with Confidence (DeepConf), set the threshold from a dedicated warm-up phase that must finish before any trace can be filtered. We introduce Peer Think with Confidence (PeerConf), which estimates the same percentile threshold from traces finishing within the run itself, a self-calibrating threshold. The threshold arms at the first finisher and is recomputed at every finisher thereafter, so filtering and consensus checks begin immediately. PeerConf also probes running traces for their current answer; a trace whose probe clears a confidence threshold and closes the model's own reasoning commits early, casts that answer as its vote, and frees its seat. Evaluated on MATH500, AIME25, and HMMT25 using DeepSeek-R1-8B and GPT-OSS-20B, PeerConf matches or exceeds self-consistency accuracy while reducing token consumption by 53.1% to 79.8%. More broadly, the calibration for confidence-filtered reasoning is already latent in a run's own finishing traces, and a model emits an identifiable commitment signal, a confident answer together with the closure of its own reasoning, when it has settled; neither needs to be paid for separately. Code and benchmarks are available at https://anonymous.4open.science/r/PeerConf.