Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
Abstract
While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it is still unclear VLMs are able to recover temporal dynamics when motion is present. To this end, we introduce the animation-to-code task, as well as a benchmark, \texttt{Animation2Code}, for evaluating temporal visual reasoning via reconstructing executable web animation code from videos. \texttt{Animation2Code} consists of 1,069 web animation videos with diverse visuals and motion patterns, augmented with corresponding HTML/CSS/JavaScript implementations. We further pair the dataset with several novel human-aligned metrics, which allow us to disentangle visual similarity from temporal alignment when comparing rendered animations to ground-truth samples. We benchmark current SOTA model performance on these new samples, and show that current models struggle with temporal consistency, even when achieving high appearance similarity, even in fine-tuning and iterative refinement scenarios. Our benchmark dashboard is available at: \url{https://anim2code-dashboard.vercel.app}.