Can Image Captions Support Concept Interpretation? A Controlled Validation Study
Abstract
Image captions provide an inexpensive, open-vocabulary source of evidence for interpreting concepts, but their suitability for evaluating concepts produced by extraction methods remains unclear. We present a controlled validation framework that separates caption-based concept interpretation into three stages: evidence availability, semantic recovery, and concept-level evaluation. Using COCO object categories and CUB fine-grained attributes as ground-truth concepts, we find that captions often contain discriminative semantic evidence, although omission is systematic and depends on visual salience. When evidence is available, its recovery depends materially on the association procedure, and increasing caption detail changes both the available evidence and its recoverability. At the concept level, balanced mixtures of genuine categories demonstrate that strong and reproducible caption associations do not establish semantic unity. Word-set coherence provides information beyond association-based signals, but we do not establish a reliable unity threshold or accept--abstain rule. Set-based precision and recall can characterize known semantic relations, yet do not reliably identify meaningful expressions or expression pairs from ranked candidates. Overall, caption-based evaluation is a promising but conditional measurement approach whose stages must be validated separately before application to extracted concepts.