“Uncertain” Is Uncertain: Prompt-Induced Abstention in VLMs Is Neither Stable Nor Human-Aligned
Abstract
Vision-language models (VLMs) increasingly output “Uncertain” responses alongside Yes/No answers, ostensibly signaling when inputs are ambiguous. We show that this signal can be highly sensitive to prompting and only partially aligned with human ambiguity. We introduce FaceTri, a trinary (Yes/No/Uncertain) benchmark of 12 fine-grained facial attributes, and evaluate four VLMs to demonstrate that prompt-induced abstention is acutely unstable: swapping the order of “Answer:” and “Reason:” shifts uncertainty rates by up to 34 percentage points, while adding a one-sentence justification changes them by up to 95 points—with direction varying unpredictably across models. To assess human alignment, three annotators provide external ambiguity labels for three attributes (Fleiss’ κ = 0.587). Model abstention correlates with human disagreement (χ 2 = 8.84, p = .003), yet is systematically overbroad: 65.5% of one model’s uncertain responses occur on images with unanimous human labels. Generated justifications are often templated and show no clear relationship to correctness. Accuracy on objectively unambiguous attributes (e.g., eyeglasses, hats) reaches approximately 0.90, versus 0.70–0.79 on the full benchmark, confirming that uncertainty is not a proxy for task difficulty. These results refute the assumption that abstention reliably indicates ambiguity. We mandate that uncertainty evaluations for subjective visual tasks report prompt protocols and results stratified by annotator agreement, enabling meaningful cross-model comparison.