AdDirector: Anchored Guidance for Generating Camera-Controllable Advertisement Videos
Abstract
Reference-to-video (R2V) generation has advanced significantly and is now a promising paradigm for controllable video synthesis. Nevertheless, in advertisement video generation, where accurate camera control and precise product presentation are essential, existing approaches still face two fundamental challenges. First, current models adhere weakly to high-level semantic camera instructions, yielding stochastic or inconsistent camera motion. Second, they struggle to preserve strict cross-frame appearance consistency, often leading to identity ambiguity artifacts such as logo drifting, texture distortion, and geometric distortion. Although certain image animation methods incorporate explicit motion conditions (e.g., trajectories or motion fields) to improve controllability, these conditions are costly to annotate, thereby constraining their practical applicability. To address these limitations, we propose a camera-aware R2V framework named AdDirector, which adaptively exploits anchored guidance from the reference image and propagates it across spatial, temporal, and scale dimensions. Specifically, AdDirector introduces Spatio-Temporal-Scale (STS) modulation, a camera-conditioned modulation mechanism that first selects multi-scale reference features, and then regulates feature injection throughout the R2V generative process spatio-temporally conditioned on camera types. This design allows our model to adaptively balance global structural coherence and fine-grained detail preservation. Furthermore, AdDirector proposes Temporal Anchoring Regularization (TAR), a boundary-conditioned feature-level regularization scheme. TAR leverages latent trajectories extracted from a base R2V model as temporal anchors, thereby reducing feature drift and improving cross-frame appearance consistency. These two modules together transform static reference conditioning into structured anchored guidance, and experiments show that our method substantially improves camera controllability, appearance consistency, and spatio-temporal coherence for advertisement video generation.