Decomposing VLM Zero-Shot Gap: Spatial, Lexical, and Format Failures on a Production Line
Abstract
Vision-language models offer a zero-shot alternative for industrial inspection, yet their viability is typically reduced to a single aggregate gap against supervised baselines, obscuring the predictable failure modes practitioners must address. We decompose this gap on an operating garment production line, evaluating ten VLMs across 2746 images and three inspection tasks. Three failure modes emerge, each isolated by a controlled intervention: spatial (attending to the wrong object), lexical (naming an attribute outside the label set), and format (producing no parseable label). Masking surrounding context improves type classification in every model-level combination tested, while its effect on color depends on scene interference; on catalog photographs without the distractor, misassignment to the interfering color falls from 62.5\% to 3.5\%. No image-level intervention removes lexical failures. Surface-form normalization recovers up to 0.426 adjusted accuracy at zero run-time cost. Format failures dominate under unconstrained prompting, and until the output is normalized they absorb lexical ones and distort the diagnosis of both, with prompt formulation spanning 0.770 adjusted accuracy units within one model and task. These measurements yield an ordered procedure: normalize, prompt, isolate. On a held-out fold the residual gaps are 0.026 for color, 0.077 for localization, and 0.189 for type. Type accuracy on curated catalog images exceeds production-line performance by up to 0.39, showing that standard benchmarks substantially overstate deployment readiness.