Agentic Video Editing from Underspecified Requests
Abstract
Recent video editing models have converged on a unified-conditioning design: a single diffusion transformer reads text, source video, reference images, and masks through one token sequence, and one set of weights covers replacement, removal, style transfer, and reference-driven insertion. The design is flexible, but it assumes that the user already provides model-ready text, identity references, and spatial targets, which real requests often omit. We present Aurora, an agentic video editing framework that pairs a tool-augmented vision-language models (VLMs) agent with a unified video diffusion transformer. The agent maps a raw user request to a structured edit plan aligned with the transformer's conditioning channels, thereby resolving textual and visual underspecification before generation. We train the agent with supervised data for executable edit planning and reference-image selection, together with preference pairs for robust tool use and instruction refinement. Moreover, we introduce AgentEdit-Bench to evaluate agent-enhanced video editing under textual and visual underspecification. Experiments on EditVerse-Bench, OpenVE-Bench, and AgentEdit-Bench show that Aurora improves over text-only baselines and its agentic capability transfers to compatible frozen video editing models. Code, data, and models will be released.