VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
Abstract
Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for text--vision tasks, typically by fine-tuning them to produce universal representations. However, their performance on video--text retrieval often remains below that of dedicated Video Foundation Models (VFMs). In this paper, we revisit the video--text retrieval capabilities of MLLMs. We first analyze the zero-shot retrieval capabilities of pretrained MLLMs and show that they already encode substantial retrieval-relevant information: combining intermediate-layer embeddings with a calibrated MLLM head yields strong performance without any training. To further exploit these capabilities, we introduce a lightweight text-based optimization strategy that requires no paired multimodal data. Our method uses dense video captions as rich textual surrogates for video semantics and short summaries as compact, query-like views, showing that a carefully designed text-alignment objective can improve multimodal alignment through the shared MLLM backbone. We backup these findings through in-depth analysis. Without visual supervision our method outperforms existing approaches, often by a substantial margin, and achieves state-of-the-art performance across common video--text retrieval benchmarks. More broadly, our results offer a new perspective on retrieval with pretrained MLLMs, suggesting that textual descriptions of visual content can serve as an effective substitute for large-scale visual supervision.