Process Rewards with Learned Reliability
Abstract
Process Reward Models (PRMs) provide step-level feedback for reasoning, but usually output only a single reward score with no indication of when it should be trusted. We propose BetaPRM, a distributional PRM that predicts both a step-level success probability and the reliability of that prediction. BetaPRM learns a Beta belief from Monte Carlo continuations through a Beta-Binomial likelihood, rather than regressing to the finite-sample success ratio as a point target. This learned reliability signal indicates when a step reward should be trusted, enabling downstream applications to distinguish reliable rewards from uncertain ones. As one application, we introduce Adaptive Computation Allocation (ACA) for PRM-guided Best-of-N reasoning. ACA uses the learned reliability signal to stop when a high-reward solution is reliable and to spend additional computation on uncertain candidate prefixes. Experiments across multiple backbones and reasoning benchmarks show that BetaPRM improves PRM-guided Best-of-N selection while preserving standard step-level error detection. ACA reduces token usage by up to 33.57\% while improving accuracy over fixed-budget Best-of-16.