Grounding in the Dark: Text-Only Training for Zero-Shot Video Temporal Localization
Abstract
Video-LLM's are typically trained on paired video-text data for temporal tasks, yet collecting precise temporal annotations and training models on large-scale video are both costly. We ask whether a video-llm for temporal grounding can instead be trained entirely on textual sequences and receive video only at inference time. This setting tests how readily the text and vision representations of a pretrained multi-modal encoder can be substituted without paired-video adaptation. We generate a 180K-example corpus of textual sequences with up to 100 time steps, train with CLIP or LLM2CLIP embeddings, and evaluate on Charades-STA and QV-Highlights without using videos from either benchmark for training. On Temporal-TextVid text, simplifying the descriptions raises CLIP mtIoU from 13.45 to 50.35, showing that the CLIP-based pipeline is highly sensitive to caption detail. When transferred to video, however, performance drops sharply to 4.37 mtIoU on Charades-STA and 5.97 on QV-Highlights, exposing a substantial modality gap. To bridge this gap, we introduce Nearest Entity Replacement (NER), which maps each video frame toward the text embedding space by representing it as the mean of its nearest text-entity embeddings. This allows the text-trained temporal model to operate on representations closer to those seen during training. We evaluate NER against four families of modality-gap baselines. NER reaches 8.79 mtIoU on Charades-STA and 12.28 on QV-Highlights, outperforming the strongest baseline by 3.25 and 2.03 points, respectively. These results show that cross-modal neighborhoods improve text-to-video benchmark transfer, however there is still a substantial gap between image and text embeddings from current multi-modal encoders making them unsuitable for direct replacement.