VideoSailor: Navigating Video Deep Research via Trajectory-to-Policy Flywheel
Abstract
Recent advances in Multi-modal Large Language Models (MLLMs) have significantly improved video understanding. However, real-world scenarios often require going beyond isolated video inputs, giving rise to the emerging paradigm of video deep research, which involves open-world reasoning through integration of external knowledge. Despite this progress, existing approaches are limited by their inability to perform effective inter-video interaction and their reliance on static, one-shot datasets that yield noisy and suboptimal supervision. To address these challenges, we propose VideoSailor, an open-world video deep research agent powered by a trajectory-to-policy flywheel that iteratively converts generated reasoning trajectories into high-quality off-policy guidance to mitigate the limitations of low-quality on-policy rollouts. For systematic evaluations, we further introduce Omni-BrowseComp, a benchmark that emphasizes omni-modal evidence aggregation, cross-video reasoning, and fine-grained temporal grounding with reasoning-relevant segment annotations. Extensive experiments demonstrate that VideoSailor significantly improves performance on complex multi-hop and cross-video reasoning tasks, establishing a scalable and effective paradigm for open-world video deep research.