HumanScore: Benchmarking Human Motions in Generated Videos
Abstract
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, there is no method to systematically assess the fidelity of human figures generated in these videos. In this paper, we present HumanScore, a systematic framework to evaluate the \textbf{quality of human figures and motions} in AI-generated videos. We introduce seven metrics spanning kinematic plausibility and biomechanical consistency. Through carefully designed prompts, we elicit a broad set of movements at varying intensities. We evaluated a total of 840 videos generated by 6 state-of-the-art models. Our framework reveals consistent gaps between visual plausibility and motion fidelity, highlights common failure modes, and ranks models from multiple quantitative and physically meaningful metrics. The proposed metrics are correlated with human evaluations and are simple to compute using standard pose estimators.