Anymotion: A Dataset, Benchmark, and Baseline for Controllable Human Motion Editing
Abstract
Human motion change is currently hindered by the lack of large-scale datasets, comprehensive benchmarks, and robust algorithms capable of handling diverse motion requirements. To address these gaps, we propose AnyMotion, a novel open-world paradigm for highly controllable and identity-consistent human motion editing across arbitrary scenes and actions. Our contributions are two-fold. First, we establish a hierarchical label taxonomy covering six major categories and 527 fine-grained classes, and curate HME-1M, a large-scale dataset containing 1.04 million editing quadruplets via systematic data balancing, cleaning, and reverse construction. Second, we introduce the Skeleton Chain-of-Thought (SCoT) framework, a two-stage pipeline consisting of motion reasoning and guided generation. In the reasoning stage, a multimodal large language model serves as a cognitive planner to derive target skeletal keypoints and MetaQueries. These outputs subsequently function as explicit structural and semantic priors, providing the necessary guidance for the diffusion transformer to perform high-fidelity motion generation. Finally, built upon our label taxonomy, we present Human Motion-Bench, the first motion-centric benchmark equipped with multi-category fine-grained evaluation metrics. Extensive experiments demonstrate that AnyMotion achieves state-of-the-art performance across diverse editing scenarios.