Split and Bridge: Multimodal Generation via Diffusion Bridging
Abstract
Visual data are naturally expressed through multiple complementary modalities (e.g., images and segmentation masks), each capturing distinct aspects of the same underlying structure. Most existing multimodal methods adopt a top-down paradigm, learning a shared latent variable through joint training from scratch, which limits scalability. We propose Split and Bridge (SnB), a bottom-up framework for multimodal generation that couples pretrained unimodal diffusion models at sampling time. Rather than learning a unified latent space, SnB generates coherent multimodal samples by following a shared diffusion trajectory up to a split point, after which modality-specific processes are guided via a bridge module. This design enables plug-and-play reuse of powerful pretrained models without joint training, making it scalable and easily adaptable. Across standard benchmarks (PolyMNIST, CelebAMask-HQ) and more complex settings (PIE-Bench++, ROCO), SnB consistently achieves a strong balance between sample quality and coherence, as we report a 33% increase in the coherence score on CelebAMask-HQ over prior top-down approaches. We show that SnB naturally extends to downstream zero-shot tasks such as interleaved co-generation, image editing, and style transfer. Our results demonstrate that bottom-up generation offers a practical and scalable alternative to top-down multimodal models.