From Structural Feedback to Prompt Policies: Learning Faithful Text-to-Image Prompt Editors
Abstract
Text-to-image prompt optimization is often treated as prompt rewriting: given a user prompt, a language model expands it with richer style, composition, and photographic details. While this paradigm can improve visual appeal, it may silently weaken the user's actual intent, especially when the prompt specifies fine-grained objects, attributes, and relations. We argue that faithful prompt optimization requires a different objective: improving perceptual quality while preserving the compositional structure that makes the prompt semantically correct. We propose \ours{}, a single-step reinforcement learning framework that transforms noisy scene-graph feedback into a learnable prompt-editing policy. Instead of freely rewriting the entire prompt, \ours{} first extracts a calibrated scene graph from the input and then performs conservative deletion, reordering, and insertion actions conditioned on both text and graph representations. This design exposes structured intent to the optimizer while restricting unnecessary prompt drift. To make object--relation feedback usable for policy learning, \ours{} replaces sparse thresholded grounding rewards with uncertainty-aware smooth rewards and estimates advantages with environment-aware grouping over multiple image samples from the stochastic text-to-image generator. Across prompt-optimization datasets and generation backbones, \ours{} improves the faithfulness--aesthetics trade-off, with particularly consistent gains in relational fidelity where generic rewriting and aesthetics-oriented optimization often fail. We further evaluate with both human preference and Gemini-VQA, showing that the improvements are not confined to the SG reward used during training. Appendix additionally compares against a high-budget iterative GPT-4o structural-feedback baseline to contextualize the efficiency of amortized policy learning. More broadly, our results suggest that structural constraints should not remain only post-hoc evaluators of generated images; when properly calibrated and stabilized, they can be internalized as training signals for reliable, intent-preserving prompt policies.