EmbodiedObject: All-in-One Object Understanding with SLAT
Abstract
Recent advances in 3D object generation have enabled the creation of high-fidelity 3D assets from 2D images. This 2D-to-3D capability offers a promising path for alleviating 3D data scarcity, which remains a key bottleneck for scalable simulation in robotics and AR/VR applications. However, existing 3D generation methods primarily emphasize visual plausibility and geometric fidelity, with limited consideration of object-level interaction and structural understanding that are critical for downstream embodied AI tasks. To bridge this gap, we introduce EmbodiedObject, a unified framework for embodied 3D object understanding built upon latent representations produced by modern 2D-to-3D generative models such as SAM3D. We first systematically analyze the part-level locality of these latent representations against existing unified 3D point encoders and show that they naturally preserve rich structural semantics suitable for object understanding. Motivated by this observation, we then design a unified DiT-style decoder that operates directly on the latent representation and supports multiple object-level tasks, including open-vocabulary 3D affordance prediction, 3D part segmentation, and 3D articulation estimation. Given a real-world 2D image containing an object of interest, EmbodiedObject extends 2D-to-3D generation beyond visual reconstruction by jointly reasoning about affordance, articulation, and part semantics. This transforms generated 3D assets into actionable object representations suitable for downstream embodied AI applications. Our source-code will be open-sourced.