Position: Telic Errors Make VLMs Unreliable Annotators in Sensitive Contexts
Abstract
VLMs now annotate sensitive content at scale, be it hate-speech memes, conflict imagery, politically charged symbols; yet, standard pipelines measure only whether a model can identify what an image depicts, not what it means to the community it concerns. We name this failure mode a telic error: the model produces a surface-accurate, perceptually complete annotation (telic: what-it-depicts) while erasing the relic meaning, i.e., the symbolic, cultural, or political significance that makes the content consequential. We evaluate six open-weight VLMs on a two-task diagnostic designed to locate the failure precisely: Task 1 (forward annotation) shows that prepending the model's own image description does not improve accuracy (-2.0pp), ruling out description-generation failure; Task 2 (label-supplied explanation) shows the same models achieve 95-100% relic identification accuracy when given the correct label and asked to explain it, ruling out knowledge absence. The 35-percentage-point gap between the two tasks is the empirical signature of a protocol artefact: models possess the cultural knowledge required for accurate annotation, but standard binary formats never ask them to deploy it. We argue that measuring only telic accuracy is not a neutral technical choice; it is a choice about whose meanings count, and that the dual-axis evaluation framework, five-type error taxonomy, and two-task diagnostic introduced here are the instrumentation needed to make that choice visible.