Long Video Instructional Editing in the Wild
Abstract
Instructional video editing has matured rapidly on short, single-shot clips, yet long videos, the dominant form of video in practice, remain largely unaddressed. Long videos exhibit failure modes that short-video editing does not stress-test: edit targets recur across temporally distant shots under varying viewpoint, scene, and scale, demanding cross-shot semantic consistency, while non-target intervals must be left unchanged, demanding accurate temporal grounding of the instruction. We present a hierarchical framework that addresses these two axes with dedicated components: a planning agent that grounds the instruction and partitions the timeline into shot-aligned segments, a long-term anchor bank that maintains a sparse set of high-quality cross-shot anchors through critic-driven iterative refinement on top of a strong image editing prior, and an anchor-conditioned video editor obtained by lightly adapting a pretrained short-video editor through LoRA, with no architectural changes. The editor is trained exclusively on short-video pairs and lifted to long videos at inference through its conditioning interface, requiring no long-video supervision. To evaluate this regime, we introduce LVEdit-Bench, the first benchmark designed for long-video instructional editing in the wild, together with a three-tier evaluation protocol that isolates per-shot fidelity, cross-shot identity preservation, and robustness to target-absent content under a structured VLM-judge protocol. Across all three settings, our method matches strong short-video editors on per-segment quality and substantially improves cross-shot consistency and grounding, with the margin widening as temporal complexity grows.