BRAVO: Bridge Matching for Autoregressive Video Generation
Abstract
Image-to-video generation aims to animate an image while preserving appearance details and producing visually coherent future frames. Autoregressive generation models this process as a sequence of next-frame predictions conditioned on the input image and text prompt. However, advanced autoregressive generators built on diffusion or flow models formulate each step as a noise-to-data process, despite the fact that the previous frame already provides an informative visual prior. This mismatch makes next-frame synthesis harder than necessary and can degrade prediction quality, since the model must recover scene layout and visual context from uninformative Gaussian noise at every step. To this end, we introduce BRAVO, a Brownian bridge matching framework for autoregressive video generation. BRAVO replaces the conventional noise-to-data path with a data-to-data bridge from the previous frame to the current one, turning next-frame synthesis into local frame-to-frame transport. The bridge state is formed by interpolating between adjacent real frames with a Brownian perturbation, so the learned velocity models the local transition from the previous frame to the next under the text prompt and causal visual history. BRAVO is trained with teacher-forcing and sampled autoregressively, where each bridge starts from the last generated frame and reuses historical messages through a causal KV-cache. Experiments on different datasets show that BRAVO surpasses the conventional noise-to-data baseline and achieves even competitive performance with substantially larger models.