Heterogeneous Parallelism for Multimodal Large Language Model Training
Abstract
Foundation model training is becoming multimodal across the stack, from natively multimodal post-training pipelines to large-scale pretraining. As multimodal coverage broadens, context windows grow, and encoder–LLM scales diverge, a single LLM-centric TP/DP/PP/CP layout increasingly bottlenecks training throughput. This coupling forces encoders to inherit LLM-driven sharding and placement choices that can introduce unnecessary communication, limit useful encoder parallelism, or constrain the LLM schedule; the mismatch is especially pronounced at long contexts, where LLM context parallelism is needed for the fused multimodal sequence but encoder inputs remain bounded. We present heterogeneous parallelism for multimodal large language model training, a training-system abstraction that lets modules in the same end-to-end graph use independent parallel layouts and rank placements, supporting colocated execution, where modules share physical GPUs under different logical grids, and non-colocated execution, where modules occupy disjoint rank sets. The key systems challenge is preserving boundary tensor semantics when adjacent modules use independent layouts: forward activations must be materialized for the destination layout, while backward gradients must be routed back to the source layout. We address this with boundary communicators, runtime primitives that implement these forward and backward layout transforms, together with scheduling extensions for both placement modes. We evaluate optimized homogeneous, colocated heterogeneous, and non-colocated heterogeneous configurations across diverse multimodal workloads and GPU scales to characterize where each placement mode helps. Across this sweep, colocated heterogeneity improves TFLOPs/s/GPU by up to 14.4\% at short context and 41.8\% at long context, while non-colocated heterogeneity improves aggregate token throughput by up to 13.0\% and TFLOPs/s/GPU by up to 9.6\%. We validate loss convergence parity against homogeneous baselines and release an open-source implementation in Megatron-LM.