See, Read, Compare: Candidate-Aware Verification for Agent Test-Time Scaling
Xinyu Ye ⋅ Yongliang Wu ⋅ Xingyu Zhu ⋅ Peng Xia ⋅ Huaijin Wu ⋅ Rui Ye ⋅ Yehui Tang ⋅ Hao Xiong ⋅ Junchi Yan ⋅ Huaxiu Yao
Abstract
Test-time scaling (TTS) improves LLM agent performance by sampling $K$ independent rollouts and selecting one with a verifier. In open-ended software engineering (SWE) agent settings, verification typically relies on code execution to obtain a direct correctness signal, but incurs substantial setup and testing cost. By contrast, execution-free verifiers offer a more scalable alternative that chooses from the submission alone. However, existing execution-free verifiers are largely submission-isolated: they fix rubrics before seeing the candidate batch, underuse the trajectory evidence already produced during agent rollouts, and rely on poorly calibrated pointwise scores for the final decision. As a result, they often miss the right candidate even when it lies within the pool. To address this gap, we propose a training-free and execution-free verifier for agent test-time scaling based on a See, Read, Compare (SRC) protocol, which treats the $K$ candidates as mutual context rather than independent inputs. The See stage combines task understanding with cross-candidate contrast to induce an instance-specific rubric. The Read stage extracts structured behavioral features, and compresses each trajectory into an evidence-grounded summary along several cognitive dimensions: localization, hypotheses, interventions, and validation. The Compare stage performs rubric-guided pairwise selection through a two-phase hybrid tournament, preserving relative judgment with only linear complexity calls. Experiments show that SRC outperforms all compared execution-free baselines across SWE-bench Verified generator pools and generalizes to heterogeneous model pools and Terminal-Bench 2.
Chat is not available.
Successful Page Load