Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems
Abstract
LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depends not only on individual generations, but also on how agents interact and evolve across rounds. We propose SPI (Sequential Probabilistic Inference), a lightweight, training-free uncertainty estimator that formulates MAS uncertainty as sequential inference over a latent system-level belief. SPI aggregates round-level agreement and generation-uncertainty signals through a filtering-style update. Across five backbones, six benchmarks, and two MAS protocols, SPI improves misclassification detection, selective prediction, and calibration over a broad set of uncertainty estimation baselines, including standard log-likelihood-based methods and MAS-specific estimators.