LongSCOP: Semantically Consistent Long Video Outpainting
Abstract
Long video out-painting is an important yet underexplored problem, while existing video out-painting methods are mainly developed for relatively short clips. However, extending video out-painting to long videos is highly challenging. A natural solution is iterative out-painting, \ie, generating videos chunk by chunk, but this introduces two major issues: unstable even crashed long-horizon generation caused by inter-chunk inconsistency and accumulated errors, and semantic drift over long temporal horizons. In this paper, we propose LongSCOP, short for Semantically Consistent long Video Out-Painting. First, we provide a fundamental training recipe with three components for smooth and stable long-horizon generation: self-forcing training, latent calibration, and future references. Built upon this stable generation foundation, a more advanced requirement is semantic consistency, for which we further develop a VLM-based agent that automatically selects the most informative reference frames from the whole video, enabling coherent subjects, scenes, and backgrounds throughout long-range generation. To support training and evaluation, we construct a large-scale, high-quality training dataset and introduce LVO-Bench, a dedicated benchmark curated by human experts for assessing long-horizon stability and semantic consistency. Extensive experiments demonstrate clear advantages of our method over existing approaches in both generation stability and semantic coherence, significantly improving the quality and reliability of long video out-painting.