HoloCode: A Code-Centric Multi-Agent Framework for Image-to-3D Scene Generation
Abstract
Image-to-3D scene generation must recover both object geometry and plausible spatial relations. A common diffusion-based pipeline, exemplified by SAM 3D, generates each object separately and assembles the results. This preserves local details, but the lack of strong scene-level constraints often causes floating objects, interpenetrations, missing contacts, or implausible relative scales. Inspired by human designers who plan a layout before placing assets, we introduce \ours{}, a code-centric multi-agent framework. Given a single image and 2D localization cues, a vision-language model (VLM)-based Arranger Agent predicts executable Blender Python layout code with object identities and 3D transforms. A Modeler Agent binds each object to geometry through asset retrieval or image-to-3D object generation, and a Refiner Agent renders the scene and edits the code from visual feedback. The code representation exposes object names, positions, rotations, and scales, making generated scenes editable and interpretable. We train the code-generating agents using Blender Python datasets built from SAGE and 3D-FRONT, then further optimize them with GRPO rewards for executability and spatial alignment. Experiments in in-domain and cross-domain settings show that \ourMethod{} improves scene-level plausibility while preserving object fidelity.