Towards Principled Efficient Rollout Allocation in Test-Time Training
Abstract
Test-time training and test-time policy optimization have emerged as effective tools for improving the robustness of LLMs under distribution shift. A common approach constructs label-free reward signals from multiple sampled rollouts using majority voting, but existing methods typically rely on a fixed, large rollout budget for every query. This fixed allocation is inefficient because many queries reach consensus well before the rollout budget is exhausted. In this work, we introduce \textbf{AdaPO}, which replaces fixed-budget majority voting with an adaptive two-stage voting procedure that moves from electing to confirming when a leader is statistically dominant, and stops once the posterior error of accepting that leader is sufficiently low. AdaPO is plug-and-play with standard policy optimization algorithms such as PPO and GRPO, and it is also compatible with supervised fine-tuning at test time. Our rigorous theoretical results show that, under the symmetric categorical noise model with known likelihood parameters, AdaPO's electing-stage trigger coincides with the classical sequential probability ratio test (SPRT), while the confirming-stage rule yields a principled posterior-threshold stopping mechanism. Across representative reasoning benchmarks, AdaPO substantially reduces compute, achieving over 40\% token savings on GPQA while maintaining competitive accuracy. The source code will be open upon acceptance at \url{https://open-upon-acceptance}.