Determine: Benchmarking Graph-Driven Agentic Movie Making
Abstract
Film production is a sequence of dependent creative operations, from character and location design through storyboarding to shot generation, yet AI film generation is evaluated end-to-end, scoring the finished video for aesthetics and narrative adherence but not which operation caused a flaw. We introduce \benchmark, which casts film generation as a directed acyclic graph of atomic operations and scores each one both in isolation, on ground-truth inputs, and in-pipeline, so defects can be attributed to their source. \benchmark pairs an end-to-end setting of 30 films specified by screenplays and shot lists with a per-operation setting that isolates each operation from upstream error, and supplies orchestrators ranging from a fixed deterministic plan to autonomous agents over a shared skill registry. Sweeping six image and six video models, we find that per-operation generation quality is an incomplete measure of a film: the weaker, cheaper agent scores higher on execution because it simplifies each shot into an easier target, no longer matching what the screenplay asked for. Furthermore, a fixed deterministic plan stays competitive without feedback, indicating that current agentic harnesses do not yet convert self-correction into a decisive advantage on cross-shot consistency.