A Unified Framework for Image-to-3D Part Generation via Variable Granularity
Abstract
Controllable 3D part generation is pivotal for digital asset creation, yet existing methods lack the flexibility to accommodate diverse user workflows requiring varying levels of control. In this paper, we introduce FlexPart, a unified image-to-3D part generation framework that supports variable-granularity geometric conditions within a single model. By seamlessly integrating heterogeneous inputs, ranging from sparse points to dense bounding boxes and masks, via gated adaptive modulation, FlexPart unlocks a powerful cross-granularity synergy. We also propose an asymmetric geometric guidance mechanism that leverages strong spatial priors to align representations across granularities, improving geometric consistency and facilitating inter-task synergy. Extensive experiments demonstrate that this synergistic approach outperforms state-of-the-art methods by over 10.5\% on Part CD, achieving high-fidelity part generation with superior geometric consistency across all prompt granularities. Code and models will be made publicly available.