Selected-Tail Reliability in Verifier-Guided Best-of-$N$ Inference
Teresa Zhang
Abstract
Today, verifier-guided best-of-$N$ inference is a common way to spend test-time compute in language modeling, code generation, and alignment: generate $N$ candidates, score them with a verifier or reward model, and deploy the highest-scoring output. This paper argues that the relevant reliability object is not average verifier accuracy but \emph{selected-tail reliability}, namely the behavior of the output chosen after optimizing over $N$ noisy public scores; we derive exact selected-rank exposure laws for best-of-$N$ selection and finite-verifier resource laws showing that selected public optimism scales at the $\sqrt{\log N/m}$ regime under $m$ units of verifier evidence. We then separate public optimism from hidden utility harm, showing that public-score inflation alone is not a harm claim and that harm or uncertifiability requires public-hidden mismatch, tail dominance, count imbalance, or an unresolved selected-tail regime. Building on these results, we introduce Tail-Certified Compute Caps, a conservative procedure that certifies candidate budgets or refuses certification using selected-tail intervals, verifier-resource penalties, mismatch envelopes, and feasibility gates. In exact-verifier/public-hidden experiments, the theory predicts selected false-positive exposure and supports resource-aware cap decisions over an all-attempted denominator, while non-code candidate tables with pre-existing public human annotations serve as secondary scoped $K\ge16$ consistency checks under proxy/model-prior/verifier-style scores across Arena55K and Stanford SHP. Overall, the paper provides a reliability framework for verifier-guided inference-time scaling: candidate budgets should be justified by evidence about the selected tail that is actually deployed, not only by average verifier validation.
Chat is not available.
Successful Page Load