Simmer: A Scalable Pretraining Recipe for Video-Text Encoders
Abstract
Most video-text encoders are pretrained on carefully curated video-text datasets where each video is paired with a single short caption. Such supervision captures only one narrow view of a video, despite videos naturally supporting many complementary descriptions of their objects, actions, and relations. We present Simmer, a scalable pretraining recipe for video-text representation learning that increases textual coverage per video and matches the training objective to this richer supervision. We first augment a video corpus with diverse synthetic captions written in many distinct textual registers, producing a large set of aspects per video. Since this many-to-one supervision is difficult to capture in a single global embedding, we pair caption diversification with a multi-vector contrastive loss that allows videos to align with multiple relevant captions. Second, we show that introducing a distinct cooldown phase over a smaller corpus of human-sourced dense captions with our stylistic augmentations further improves performance. Our resulting encoders achieve state-of-the-art performance on five video-text encoding tasks across three model scales, outperforming prior works while using 2-60x fewer videos. Our results show that broader caption coverage and multi-aspect alignment can pave the way towards new scaling laws in video-text modeling.