LogicDirector: Enforcing Temporal Composition in Text-to-Video Generation
Abstract
Accurately rendering temporal composition is fundamental to text-to-video (T2V) generation. While recent efforts incorporate temporal structure for long-horizon coherence, the intrinsic temporal understanding capability of T2V models remains underexplored. We observe that current models still struggle to respect canonical operators such as before and after, usually accompanied by missing entities and incorrect action binding. In this work, we propose LogicDirector, a test-time guidance framework that enforces temporal composition through executable logical specifications. By translating text prompts into first-order logic constraints over attention-derived evidence, LogicDirector constructs a neuro-symbolic verifier to evaluate noisy latent states during diffusion sampling for entity grounding, action binding, and temporal adherence. To enforce these constraints without expensive gradient-based optimization, we introduce a gradient-free Best-of-n latent search strategy, which preserves the model's native diffusion dynamics while steering logical satisfaction. To systematically study operator-level temporal composition, we curate TempBench, a diagnostic benchmark spanning four canonical relations, together with a hierarchical evaluation protocol that disentangles entity existence, event realization, and temporal correctness. Experiments on CogVideoX and Wan 2.1 demonstrate that our method significantly improves temporal composition while maintaining visual fidelity.