Block Parallelism for Efficient Distributed Long-Context Diffusion Language Model Training
Tarun Suresh ⋅ Pranshu Chaturvedi ⋅ Hangoo Kang ⋅ Parth Shroff ⋅ Ishan Khare ⋅ Hermann Kumbong ⋅ Azalia Mirhoseini
Abstract
Block diffusion language models aim to combine the generation quality of autoregressive models with the efficiency of diffusion-style parallel generation, but training them over long contexts is limited by distributed attention and activation memory. Conventional context parallelism (CP) partitions the combined clean-plus-corrupted sequence by token position, so it communicates both shared clean K/V and block-specific corrupted K/V during forward and their gradients during backward. Our insight is that we can eliminate this block-specific communication by exploiting the BDLM objective's per-block loss separability to expose target-block computation as a new distributed parallelism dimension. We introduce block parallelism (BP), which assigns each complete corrupted-block computation to one rank while keeping the block's corrupted K/V and gradients local. BP alone replicates long, overlapping clean prefixes. We therefore introduce context-sharded block parallelism (CSBP), which applies CP to the shared clean sequence over the same ranks while retaining complete corrupted-block computations locally. CSBP restricts cross-rank attention communication to shared clean K/V and their gradients without replicating clean prefixes. On two nodes with 16 H200 GPUs, CSBP improves throughput at 256K context by $\textbf{1.18--1.57$\boldsymbol{\times}$}$ for supervised fine-tuning and $\textbf{1.37--1.40$\boldsymbol{\times}$}$ for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. These gains are measured against the highest-throughput tested combinations of existing parallelism strategies and batch size. CSBP's advantage grows with context length: from 64K to 512K, speedup increases from $\textbf{1.10$\boldsymbol{\times}$ to 1.19$\boldsymbol{\times}$}$ on Nemotron 14B and from $\textbf{1.26$\boldsymbol{\times}$ to 1.68$\boldsymbol{\times}$}$ on DiffusionGemma 26B-A4B.
Chat is not available.
Successful Page Load