Beyond Copy-Paste: How Well Do Subject-Driven Video Models Understand Their Subjects?
Abstract
Subject-driven video generation aims to produce videos that faithfully incorporate user-provided reference subjects. Current evaluation relies on free generation followed by reference-similarity scoring, which rewards a fundamental shortcut: models can obtain high scores by reproducing visible cues from the references while avoiding transformations that would test whether the subject is truly preserved beyond the conditioning input. We design anchored evaluation: controlled settings in which generation is anchored to held-out targets of the same subject rather than left unconstrained. Each probe withholds target content from the model's conditioning, so success requires using the references to infer content not directly visible, rather than simply copy-pasting what was given. We instantiate two probes. (1) Viewpoint control specifies target geometry via depth-based warping, leaving occluded regions that the model must complete using subject appearance from the references. (2) Degradation--recovery corrupts held-out videos showing the subject in real-world contexts that differ from the reference images, then recovers them conditioned on those references; recovery requires cross-context generalization of identity and inference of state dynamics. Both probes require only lightweight modifications to the inference loop of flow-matching models, need no retraining, and generalize across architectures. Evaluation across seven diverse, strong models reveals that rankings shift substantially: VACE-14B ranks first conventionally but Phantom-14B, ranked last, leads under anchored evaluation. Moreover, conventional scores strongly correlate with the improvement from providing ground-truth references (Pearson r = 0.91), suggesting they largely measure copying ability. Fine-grained analysis further reveals distinct failure modes: copy-paste models drift at unseen viewpoints, and no model adaptively balances copying and generalization.