ELMA: Benchmarking Anaphoric Compositions for Long-Term Text-to-Motion Generation
Abstract
Text-guided long-term human motion generation is essential for applications such as continuously operating digital avatars. Despite recent progress on short clips, existing models struggle with multi-segment instructions that involve anaphora, where constraints established in earlier segments must be preserved over time. In this paper, we introduce a novel task of long-term motion generation from such anaphoric compositions. To mitigate this, we present ELMA, an episodic long-term motion dataset with dense segment-level annotations that preserve cross-segment dependencies. Built via a fully automated collection and filtering pipeline, ELMA extracts high-quality, continuous 30-second clips from in-the-wild videos, capturing the episodic narratives necessary for long-range state modeling. Leveraging this dataset, we fine-tune state-of-the-art motion generation models to incorporate long-range context. Extensive experiments reveal fundamental shortcomings in current models when handling context-dependent instructions. We demonstrate that ELMA enables significant improvements in both global physical realism and semantic alignment, setting a robust new baseline for long-term motion synthesis.