ApertureAttn: Native 4K Video Generation with Image-Only Supervision
Abstract
Recent progress in video generation has been remarkable, yet state-of-the-art models remain largely confined to 720p, well short of the 4K resolution of modern displays. Scaling video generation to higher resolutions is constrained by the scarcity of high-resolution video data and the prohibitive cost of large-scale training. Beyond data and compute, higher resolutions place substantially greater demands on spatial modeling, particularly for fine-grained details, complex compositions, and local structures. Such high-resolution spatial supervision, however, is abundantly available in high-resolution images, which are easier to collect and train on at scale. This naturally raises a question: can image-only supervision unlock native 4K video generation? To answer this, we propose ApertureAttn, an inverted attention pyramid that concentrates adaptation on high-resolution fine-grained spatial modeling using only single-frame high-resolution images, without any video training data. We further introduce window-scaled extrapolation and proximity-guided temporal attention to enhance spatiotemporal consistency. Extensive experiments show that our method enables native 4K video generation with strong visual fidelity and temporal coherence, while requiring only image supervision and minimal training cost. Despite this lightweight training setup, it remains competitive with state-of-the-art high-resolution video generation methods trained directly on high-resolution video data at substantially higher training cost.