The Design Space of Tri-Modal Masked Diffusion Models
Louis Bethune ⋅ Victor Guilherme Turrisi da Costa ⋅ Bruno Mlodozeniec ⋅ Pau Rodriguez ⋅ Lokesh Boominathan ⋅ Nikhil Bhendawade ⋅ Amitis Shidani ⋅ Joris Pelemans ⋅ Theo X. Olausson ⋅ R Devon Hjelm ⋅ Paul Dixon ⋅ Joao Monteiro ⋅ Pierre Ablin ⋅ Vishnu Banna ⋅ Arno Blaas ⋅ Nick Henderson ⋅ Kari Noriy ⋅ Dan Busbridge ⋅ Marco Cuturi ⋅ Joshua Susskind ⋅ Irina Belousova ⋅ Luca Zappella ⋅ Russell Webb ⋅ Jason Ramapuram
Abstract
Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent multimodal work either finetuning unimodal diffusion bases or distilling autoregressive backbones for bi-modal generation. Diverging from these approaches, we introduce the first tri-modal \gls{mdm} \emph{pretrained from scratch} on text, image-text, and audio-text data, where all generation tasks are learned simultaneously within a single model. We systematically study the design space governing stability and efficiency at scale: we conduct the first empirical characterization of the critical batch size $B_{\text{crit}}$ under the SDE reparameterization of AdamW, and introduce a drift--horizon interpolation parameter $\gamma$ that balances gradient noise reduction and optimization horizon when scaling the token budget. We derive multimodal scaling laws and ablate modality mixing ratios, noise schedules, and anti-masking. Lastly, we pretrain a 3B model on 6.4T tokens, achieving competitive results in text generation, text-to-image, and text-to-speech.
Chat is not available.
Successful Page Load