Training Agent to Scale Inference-Time Reasoning
Abstract
Inference-time scaling provides a promising path for improving agentic reasoning performance, and one widely adopted form is sample-then-select: first sample a pool of candidate reasoning trajectories, then select from the resulting candidates. While this keeps scaling lightweight without updating the policy, it exposes a training-deployment mismatch. Standard trajectory-wise Reinforcement Learning (RL) objectives optimize normalized signals for individual trajectories, whereas choosing from a large candidate pool with the policy's own selection signal is inherently set-conditional: it requires calibrated ranking and confidence aggregation across correlated trajectories. To address this challenge, we propose AutoPortfolio, a set-conditional policy alignment framework for efficient and effective sample-then-select scaling. AutoPortfolio builds a tractable environment-informed target on the sampled candidates, and uses its induced Plackett-Luce (PL) ranking to calibrate the policy which trajectories to prioritize within that set. It trains the policy with a balanced PL-motivated objective that reallocates probability mass among competing trajectories, preserving reward-aligned alternatives while sharpening away from low-quality distractors. Aligned with our policy training, our lightweight inference-time Mass Aggregation scaling adaptively perceives confidence from candidate trajectories and selects the most promising answer(s). Theoretically, we show that reducing our set-conditional loss narrows the discrepancy between policy-induced and environment-informed PL rankings on the sampled set. Empirically, across 12 challenging reasoning benchmarks, AutoPortfolio improves over strong agentic RL baselines, with the clearest gains in larger candidate pools where the set-level selection is crucial.