A Training-Free Video Moment Retrieval Framework via Selection from Multiple Visual Prompts
Abstract
Video Moment Retrieval (VMR) aims to localize a temporal segment in a video based on a natural language query. While Multi-modal Large Language Models (MLLMs) are increasingly adopted for VMR due to their strong video understanding capability, they still face several limitations: (1) imprecise perception of event boundaries, (2) rigid dependency on precise object identification, and (3) inability to handle temporarily interrupted events. To address these, we propose three visual prompts: Event Segmentation (ES) prompt, Object Detection (OD) prompt, and Keyframe Marking (KM) prompt. Each prompt improves performance when applied individually, yet direct fusion degrades results because mixing prompts overcomplicates video content and inaccurate labels introduce noise. Although certain visual prompts may mislead the model, at least one of the three is likely to yield a positive effect. Even in the worst case where all three perform poorly, the prediction can fall back to the baseline. Motivated by this, we introduce a selection mechanism based on Yes/No Probability-Difference score to choose the most reliable prediction among the results produced by the three visual prompts and the baseline. We evaluate our method in a training-free setting across three datasets and four models, and obtain state-of-the-art performance.