When One Word Changes an Image Judgment: A Valence Asymmetry in Vision-Language Models
Syed Islam ⋅ Sneheel Sarangi ⋅ Charlotte Li ⋅ Arnav Gowda ⋅ Ruizhe Li
Abstract
Vision-language systems increasingly read images of people alongside text that a user supplies, and the two can disagree. When they do, a deployed pipeline inherits whichever cue the model follows, so an imbalance between them is a grounding failure with a direction. We ask whether positive and negative wording pull an image's emotion judgment equally hard. We pair EMOTIC photographs with one-sentence contexts and measure how a conflicting context shifts that judgment away from a neutral-context baseline. In the most controlled condition we hold the photograph and the described event fixed and flip a single valence word (won$\leftrightarrow$lost, wonderful$\leftrightarrow$devastating). Comparing the two conflict directions on those pairs, the mirror contrast is $+0.496$ in the reported frame and question, with a sentence-resampled interval excluding zero; its sign holds across seven wordings but its magnitude does not. Within positive images, where the negative sentence conflicts and the positive one agrees, negative wording moves Qwen3-VL-8B's judgment more than four times farther on all six pairs (within-item contrast $+1.148$ $[+0.94,+1.34]$). Run without an image, the same six pairs show no detectable difference, though six pairs cannot establish equivalence. A set of six unrelated events points the same way but does not survive resampling the events. In Gemma-3-4B, a different model where a text-trained valence probe is available, replacing downstream states at text positions restores 88–93% of the context difference, and a direction estimated only from valenced text still moves the answer under conflict. Across four models the conclusion depends on the summary score and on how multi-token labels are scored: scoring only a label's first piece manufactures a null in one model and reverses the categorical ordering in another. Negative text can therefore exert disproportionate influence on image judgments. Evaluations should test what each cue says, not only which modality carries it.
Chat is not available.
Successful Page Load