CAST: Compositional Constraint Satisfaction via Attention Steering in Text-to-Image Generation
Arushi Jain ⋅ Shubham Paliwal ⋅ Monika Sharma ⋅ Vikram Jamwal
Abstract
Recent text-to-image (T2I) models generate visually realistic images but often fail to satisfy compositional constraints expressed in natural language, including object-attribute associations, requested object counts and spatial relationships. We present CAST, a training-free inference-time framework that improves compositional constraint satisfaction through attention steering. Given structured constraints extracted from a prompt, CAST formulates task-specific attention objectives that encourage correct attribute grounding, counting accuracy and spatial consistency. To support diverse generative architectures, we introduce architecture-aware interventions for both U-Net cross-attention and transformer joint-attention models while retaining a shared optimization formulation. We evaluate Stable Diffusion 2.1, SDXL, Stable Diffusion 3, FLUX-Schnell, and PixArt-$\alpha$ on $900$ prompts from T2I-CompBench++ spanning attribute binding, counting and spatial-relation tasks. CAST consistently improves compositional fidelity across all architectures and provides additional gains when combined with verifier-guided reranking. We further introduce an order-controlled VLM-based evaluation protocol that measures whether benchmark improvements correspond to perceptually recognizable gains in prompt faithfulness. Our results demonstrate that intermediate attention representations provide an effective and interpretable mechanism for enforcing structured prompt constraints without modifying model parameters.
Chat is not available.
Successful Page Load