SAMoR: Motion Modelling for Articulated Objects of Any Skeleton and Topology
Yuhao Zhang ⋅ Gerard Pons-Moll ⋅ Tolga Birdal
Abstract
Modeling motion for articulated objects of arbitrary skeleton topology remains difficult: existing motion generators target a fixed human skeleton, and prior adaptations either fail to share a vocabulary across rigs or discard motion detail through global pooling. Our key observation is that while joint-level motion does not correspond cleanly across species, motion of functional joint groups does---a human arm, a wolf foreleg, and a bird wing share semantic motion structure despite differences in joint count and connectivity, a correspondence that joint names (e.g., ``forearm'', ``wing\_L1'') partially expose even when topology does not. We introduce SAMoR (Skeleton-Aware Motion Representation for Articulated Objects), a cross-topology motion representation that encodes each motion segment as a small fixed number ($K{=}8$) of part tokens shared across arbitrary skeletons. SAMoR takes three input signals---per-joint motion features, kinematic graph structure, and joint-name embeddings---and processes them with a graph-transformer encoder, then compresses the resulting heterogeneous per-joint features into part-level tokens via cross-attention pooling and residual vector quantization, yielding a discrete motion codebook shared across rigs. To prevent the part queries from collapsing into redundant global representations, we introduce a topology-agnostic attention supervision loss, combined with random joint-name dropout to prevent over-reliance on text labels; together these encourage the part tokens to cluster joints into functional groups from names, structure, and motion jointly. We curate a unified heterogeneous motion corpus from HumanML3D, Truebones Zoo, and animated Objaverse-XL assets, and evaluate SAMoR on held-out characters with unseen skeletons. The resulting representation supports accurate reconstruction and cross-topology motion transfer, and further enables text-conditioned generation and localized part-wise editing through a MaskGIT token generator. SAMoR reaches $2.75\!\times\!10^{-2}$ normalized MPJPE on cross-topology reconstruction---$5.8\!\times$ below the strongest adapted variable-$J$ tokenizer baseline---and remains competitive with fixed-skeleton specialists on HumanML3D for both VQ-VAE reconstruction and text-to-motion generation.
Chat is not available.
Successful Page Load