Intrinsic-Preserving Schrödinger Bridge for Direct Part-Aware 3D Generation
Abstract
Controllable 3D generation is a fundamental challenge in computer vision. Existing classifier-free guidance methods often yield suboptimal alignment with the given conditions, while Schr"{o}dinger Bridge approaches overlook intrinsic geometric discrepancies across multimodal manifolds, leading off-manifold trajectories and semantic misalignment. To address these issues, we propose Intrinsic-Preserving Schr\"{o}dinger Bridge (IPSB), a novel framework for direct part-aware 3D generation that leverages cross-modal intrinsic consistency. Specifically, we introduce Hierarchical Part-level Cross-modal Encoder that constructs hierarchical part-level representations through probability-based soft coverage, explicitly constraining graph topology to establish semantically consistent part graphs across modalities. Additionally, our IPSB incorporates asymmetric GW-inspired relational regularization, which strictly preserves the intrinsic semantic structure in the conditional manifold while allowing sufficient geometric flexibility in the 3D target manifold. Based on our IPSB framework, we devise Compositionally Controllable 3D Generation based on part-wise intervention inference, enabling independent replacement or resampling of target part nodes while fixing the trajectories of remaining components, thus achieving part-level controllable editing without full regeneration. Extensive experiments demonstrate that our superiority, efficiency and generality.