Anytime-Valid PAC-Bayes Certificates for Adaptive Test-Time Scaling
Xiaoyu Li ⋅ Jiangxuan Long
Abstract
Test-time scaling systems often choose how much to sample, which prompt to use, and which verifier or aggregation rule to trust after repeatedly inspecting a calibration set. Standard validation error bars are not valid after this adaptive selection. We model a test-time scaling recipe as a randomized predictor with a random cost and prove a time-uniform PAC-Bayes certificate that holds simultaneously over all calibration times, all posterior mixtures over recipes, and all data-dependent stopping rules. Thus the same event covers the recipe search that precedes deployment, not merely the fixed recipe finally reported. For a finite grid of budgets and verifiers, the penalty is the familiar $\log M/n$, while non-uniform priors and localized posterior mixtures give sharper certificates for structured searches. We specialize the result to self-consistency and show a complementary margin law: plurality voting improves exponentially in the sample budget only on examples where the correct answer is the generator's unique mode; otherwise more compute can provably amplify errors. The same margin analysis yields an oracle water-filling law for compute allocation. A synthetic audit quantifies the optimism caused by adaptive budget selection, and a 192-example GSM8K audit using gpt-4.1-mini generations illustrates risk-cost reporting for a real inference pipeline. The results give a lightweight, distribution-free way to report risk and compute claims for adaptive inference systems.
Chat is not available.
Successful Page Load