Text-to-Image Faithfulness Metrics Track Intent, Not Realization
Chayan Malkari ⋅ Pranav S Reddi ⋅ Akhil Kuchibhotla ⋅ Grace Zheng ⋅ Ane Zuniga
Abstract
Text-to-image diffusion models often fail to render what a prompt asks for, and a growing set of faithfulness metrics is used to detect these failures. We show that three widely used metric families (Causal Relevance, cross-attention, VQAScore) agree with human judgment at high rates for the wrong reason: they behave like a predictor that ignores the image and relies on the prompt. This pattern was invisible in prior validation because the baselines used (chance, positional heuristics, attention scrambling) pass whether a metric reads the image or the prompt. We audit cross-attention across two diffusion architectures (UNet via SDXL, MMDiT via FLUX.1-dev), VQAScore on MMDiT, and Causal Relevance on multi-step generation chains (SD1.5). Our analysis includes an exhaustive sweep over 1,824 head-layer-timestep-quartile combinations (456 head-layer cells $\times$ 4 denoising quartiles) in MMDiT enabled by a custom attention extraction pipeline. No aggregation of cross-attention rescues it, and Causal Relevance is confounded by a compositional lock in the generation pipeline. Current faithfulness metrics track what the prompt requested rather than what the image rendered on the cases where a faithfulness metric is deployed to detect the difference.
Chat is not available.
Successful Page Load