Exploring MLLM-Diffusion Information Transfer with MetaCanvas
Abstract
Multimodal learning has advanced visual understanding through powerful multimodal LLMs (MLLMs). In visual generation, however, these models are often used only as global text or context encoders for diffusion generators, limiting their ability to provide structured spatial and temporal guidance. This creates an interface gap: MLLMs can parse complex layouts, attributes, and knowledge-intensive scenes, yet current generation pipelines often struggle to transfer such understanding into images and videos with precise, controllable structure. We propose MetaCanvas, a lightweight framework that organizes MLLM representations into spatially indexed canvas tokens and interfaces them with diffusion generators through patch-wise residual fusion. This provides structured spatial and spatiotemporal conditioning rather than relying only on a single global embedding or 1D query sequence. We implement MetaCanvas on three diffusion backbones and evaluate it across six generation and editing tasks requiring precise layouts, robust attribute binding, and fine-grained multimodal control. Under matched settings, MetaCanvas consistently improves over global-conditioning baselines, narrows the gap to specialized closed-source visual generation and editing systems, and unlocks new editing and in-context capabilities for video diffusion backbones. These results suggest that spatially aligned MLLM-derived canvas tokens provide a promising latent interface for transferring multimodal understanding into diffusion-based generation.