FlowTrack: Controlling Edit-Signal Execution in Inversion-Free Flow Video Editing
Abstract
Text-guided video editing requires applying a semantic change while keeping source motion, layout, and unchanged content stable. Recent inversion-free flow editors derive effective edit signals from target-source velocity differences, yet typically inject these signals directly at each solver step. In neural video sampling, however, this direct execution rule overlooks an important distinction: the velocity difference is an online, state-dependent observation rather than a fixed endpoint-displacement command. As the target branch evolves under previous injections, direct injection may accumulate temporal fluctuations and off-target perturbations, producing motion drift, identity flicker, or unintended changes in weak-response regions. We propose FlowTrack, a causal execution controller for inversion-free flow video editing. FlowTrack leaves the pretrained generator and raw edit signal unchanged, but executes the signal through a carried displacement-like direction and response-aware attenuation. We derive this update as the closed-form solution of a local execution objective that balances current signal tracking, temporal carrying, and weak-response suppression. On the full FiVE-Bench protocol, FlowTrack improves motion fidelity, unedited-region fidelity, and structure consistency over strong training-free diffusion- and flow-based editors while maintaining competitive target-prompt alignment. Code is available at https://anonymous.4open.science/r/FlowTrack.