Hider–Seeker Self-Play: Geometry-Verifiable Process Rewards for Long-Horizon Visual Search
Abstract
Long-horizon visual search requires a vision-language policy to execute multi-step spatial actions over high-resolution images, where performance depends on the entire trajectory rather than any single decision. Existing supervision breaks at this horizon: outcome-only RLVR collapses many geometric decisions into a single delayed signal and discards the per-step geometry the environment already computes, while LLM-judged step rewards recover density at the cost of verifiability, scale linearly with trajectory length, and conflate distinct failure modes. We propose HSSP, a self-play framework in which a Hider adaptively weights challenging skeletons from a fixed trajectory pool and a Seeker learns to solve them through grounded spatial actions. Instead of relying on external judges, HSSP supervises every step with a geometry-verifiable process reward built from target-directed progress and exploration coverage, with trajectory length and spatial redundancy enforced as separate Constrained Markov Decision Process (CMDP) constraints. We pair HSSP with V-Trace, a trajectory-level companion dataset that augments skeleton seeds from established benchmarks with target-region annotations and explicit search-regime labels, ready for any trajectory-level training framework. The reward design admits closed-form guarantees on backtracking, find-versus-explore separation, and policy invariance, providing theoretical backing for the observed gains. Extensive empirical results on eleven challenging multimodal benchmarks, supported by controlled ablations and qualitative case studies, show that HSSP consistently outperforms advanced baselines and transfers zero-shot to OOD tasks.