Training-Free Moment Retrieval in Unseen Hour-Long Videos
Abstract
Long Video Moment Retrieval (LVMR) aims to localize target moments in long videos from natural-language queries. While recent training-free methods have shown that temporal grounding is possible without task-specific supervision, long-video moment retrieval remains difficult due to Moment Aliasing: temporally distant regions can exhibit similarly strong responses to the same query because of repeated states, visually similar contexts, and partially matching distractors. Consequently, selecting a span from a single strong response often leads to partial matches, imprecise temporal boundaries, or missed intervals. We propose Point-to-Span (P2S), a training-free framework that addresses this problem through evidence-guided point-to-span conversion. P2S first generates reliable initial proposals from sparse retrieval evidence via Adaptive Span Generation, and then refines them with Auxiliary Evidence Elicitation and Evidence-Guided Span Refinement. Experiments on MAD and MomentSeeker show that P2S consistently outperforms prior training-free baselines in the long-video setting, with especially strong gains under stricter localization criteria, indicating improved span disambiguation rather than merely coarse temporal retrieval. These results show the importance of explicitly addressing long-video ambiguity for robust training-free retrieval.