UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
Abstract
We introduce spatially grounded contextual image generation, a new controllable image generation task that reframes the conditioning paradigm. Instead of supplying a reference image and a global text prompt through two separate encoders (vision and language), UniVL is trained to bind semantics to spatial locations directly from a single unified visual input, in which the textual instruction is rendered onto the spatial mask, removing the need for a standalone text encoder at inference. This enables contextual image generation, which follows user’s specified what should appear where instructions, as well as waiving the need of text encoder to save computation significantly. For the task, we propose a framework in which the UniVL encoder—adapted from an optical-character-recognition-pretrained backbone—reads the unified condition optically, producing a UniVL embedding fVIL that fuses visual and semantic intents to spatial locations, packed as a single token sequence. A two-stage pipeline aligns UniVL in VAE embedding space and then conditions a pretrained diffusion backbone entirely on UniVL embeddings, eliminating the standalone text encoder (e.g., T5). The reframing is deliberately minimalist for text, but the empirical payoff is large. On UniVL-ImgGen, a benchmark of 477K mask-annotated images that we construct to support training and evaluation, UniVL achieves superior image quality over text-prompted baselines (FID: 14 → 11, PSNR: 16 → 20) while eliminating the text encoder entirely, reducing inference TFLOPs by up to 52% and runtime by up to 44%. Additional ablation studies verify components of different parts of the proposed method, paving way for efficient and spatially grounded image generation with unified conditioning paradigm.