AgenticVBench: Can AI Agents Complete Real-World Video Production Tasks?
Abstract
Video production workflows offer a rich and demanding testbed for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this end, we introduce AgenticVBench, a benchmark of 100 tasks across 4 task families spanning the real world post-production workflow, constructed from the representative work of 20 industry experts averaging 6 years of experience across video production contexts. Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics, both authored under expert protocol. We evaluate frontier vision-language models (VLMs) with vendor-native and open-source harnesses. The best evaluated model achieves a score below 30\%, substantially lower than frontier performance on existing multimodal benchmarks. We further find that the choice of harness substantially affects model behavior, including but not limited to scores and failure modes. AgenticVBench provides a foundation for understanding and improving both models and harnesses for agentic video production.