EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
Abstract
Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use independently generated prompt sets with limited entity coverage and simple consistency metrics, making standardized comparison across methods difficult. We introduce EntityBench, a benchmark consisting of 140 episodes (2,491 shots) derived from real narrative media, with explicit per-shot entity schedules tracking characters, objects, and locations simultaneously across easy, medium, and hard difficulty tiers of up to 50 shots, 13 cross-shot characters, 8 cross-shot locations, 22 cross-shot objects, and recurrence gaps spanning up to 48 shots. EntityBench pairs the dataset with a three-pillar evaluation framework that disentangles intra-shot visual quality, prompt-following alignment, and cross-shot entity consistency. Cross-shot consistency, the central pillar, evaluates each recurring entity through both embedding similarity and LLM per-criterion judgment across entity-type-specific dimensions, with a fidelity gate that admits accurate entity appearance. To establish baselines, we propose EntityMem, a memory-augmented generation system that plans and stores verified per-entity visual references in a persistent memory bank before generation begins, enabling the video backbone to retrieve each entity's appearance across shots. Experiments on EntityBench show that cross-shot entity consistency degrades sharply with recurrence distance in existing methods, and that explicit per-entity memory yields the highest character fidelity (Cohen's d = +2.33) and presence among methods evaluated.