UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image
Abstract
Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and motion from sparse observations remains challenging. Existing approaches remain largely constrained by lack of supervised data or lack priors needed to reliably recover articulation, hidden geometry, and internal object structure. We present the first hierarchical agentic approach for articulated 3D object reconstruction from text or image inputs, combining hierarchical, debate-based reasoning with a video generative prior for articulation modeling. High-level agents reason about object semantics and motion using knowledge from vision-language and video models, while low-level agents estimate articulation parameters and interaction points; together, they engage in structured debate to resolve ambiguities in structure and motion. To enable reliable generation, including interior object structures, from only text or image inputs, we introduce a video generative model prior that not only synthesizes plausible object motions, but also expose occluded interiors and geometry that cannot be inferred from a single static view. By combining agentic reasoning with a video generative prior, our approach jointly infers articulation and reconstructs complete 3D articulated objects, producing high-fidelity geometry, internal structure, and motion-consistent states beyond directly observed surfaces.