HorizonComposer: Spatiotemporally Consistent Driving Video Editing with Enriched Traffic Semantics
Abstract
Instruction-guided video generative models offer a scalable solution for stress-testing autonomous driving systems by simulating diverse environmental conditions. However, driving scenes operate within a dense Operational Design Domain (ODD) governed by strict traffic semantics. When applied to global driving-condition edits, general-purpose instruction-guided video editors frequently corrupts these semantics, leading to structural hallucinations and ``temporal shock''. This work strictly distills and protects traffic semantics during video editing at three levels. First, we construct a hybrid dataset by combining temporally aligned real-world videos with synthetic edits, balancing foundational scene structure with unambiguous appearance signals. Second, during training, we cast model alignment as offline reinforcement learning via advantage conditioning. We extract dense, pixel-level rewards from the hybrid data to serve as a localized signal and actively steer generation toward structurally reliable pixel states. Finally, our model mitigates temporal shock by introducing thinking frames, a transition mechanism that seamlessly reconciles static image conditioning with dynamic video context. Extensive experiments demonstrate that our method prevents the corruption of multi-agent dynamics, establishing state-of-the-art spatiotemporal consistency and yielding an overwhelming user preference (88 % relative improvement in content preservation and 75 % in temporal consistency) and yielding a significant 5 % gain in downstream road line segmentation. We provide more video results in the supplementary material.