Scene Graphs for Preserving Compositional Semantics in Text-to-Image Reinforcement Learning
Utku Bahçıvanoğlu ⋅ Gregorio Ceria ⋅ Federico Pezzoli ⋅ Luca Biggio ⋅ Chiara Plizzari ⋅ Riccardo Barbano
Abstract
Reinforcement learning fine-tuning of text-to-image diffusion models for compositional generation typically scores objects, attributes, and relations separately, and aggregates these scores into a single image-level reward without preserving their dependencies. Scene graphs make precisely this structure explicit by binding attributes to object nodes and relations to their endpoints. We first introduce SG-Reward, a scalar reward that organizes these scores according to the dependencies encoded by a prompt scene graph while keeping the underlying scorers fixed. Building on this, we introduce SG-GRPO, which carries the same graph structure into Group Relative Policy Optimization (GRPO) advantage construction. On SD3.5-M, SG-Reward outperforms a graph-free reward that combines the same factor scores on GenEval 2, ConceptMix, and DSG-1k. SG-GRPO achieves the strongest numerical performance, reaching a Soft-TIFA GM of $28.3$ on GenEval 2, compared with $18.2$ for the base model. These results suggest that scene-graph structure provides an effective inductive bias for post-training with Flow-GRPO.
Chat is not available.
Successful Page Load