Novelty and Correctness: Task-Aligned Uncertainty Estimation for Reliable Vision-Language Models
Abstract
Task-specific finetuning of vision-language models (VLMs) can produce schema-valid outputs even as deployment inputs drift and accuracy degrades. We study task-aligned uncertainty estimation, which places estimators according to the task's failure modes. We evaluate a task-specific finetune of Qwen3.5-9B on 9,500 shifted inputs spanning six corruption families and a held-out rendering shift. Decoder-side signals detect incorrect outputs well (0.95 AUROC) but are weakly sensitive to visual shift (0.62 AUROC). Conversely, a novelty estimator over vision-encoder activations detects visual shift strongly (0.92 AUROC) but discriminates correctness weakly (0.61 AUROC). The novelty signal increases monotonically with shift severity and becomes informative before task accuracy degrades. Token-exact labels further show that softmax confidence drops sharply at the first incorrect token and recovers within a few tokens. These results indicate that constrained VLM tasks benefit from separate, task-aligned estimators for input novelty and prediction correctness.