FacEDiT: Talking Head Video Editing via Facial Motion Infilling
Abstract
Editing a local segment in a talking head video without reshooting the entire scene remains challenging and underexplored. Given local edits, such as insertion, deletion, or substitution, talking head video editing must synthesize a replacement segment while leaving the rest unchanged. This requires variable-duration local rewriting while preserving identity, unedited regions, and boundary continuity, which standard speech-driven generation and lip synchronization do not directly address. Moreover, direct supervision is infeasible, as it requires paired videos of the same person and scene differing only in a local spoken segment, which does not exist in the real world. We instead formulate talking head video editing as facial motion infilling, a self-supervised pretext task that recovers masked facial motion from speech and surrounding motion context in ordinary video--speech pairs. The key insight is that local video editing can be simulated during training by masking a motion span and reconstructing it from speech and visible motion context. Based on this formulation, we introduce FacEDiT, a mask-controlled talking head model with local temporal attention bias and temporal smoothness regularization for improved lip--speech alignment and transition continuity. We also introduce FacEDiTBench, the first benchmark for talking head video editing, covering diverse edit types and lengths with dedicated evaluation metrics. Extensive experiments show that FacEDiT produces accurate, speech-aligned edits with strong identity preservation and seamless boundary transitions. Beyond editing, the same facial motion infilling model extends to portrait animation and lip synchronization by simply changing the mask pattern, establishing FacEDiT as a unified framework for talking head video editing and generation. \textit{We will release the code and data upon acceptance.