ViPADIA: A Lightweight Robust Visual Prompting Machine with applications to Industrial Anomaly Detection and Reasoning
Abstract
Large Vision-Language Models (VLMs) have shown strong general-purpose visual reasoning capabilities, yet their effectiveness in specialized industrial domains remains limited. In this work, we study industrial anomaly analysis across multiple domains and evaluate whether visual prompting can improve VLM performance without expensive fine-tuning. We introduce WTQA and CCQA, two high-quality, hand-labeled visual question answering benchmarks comprising 781 wind-turbine and 900 concrete samples, respectively, spanning anomaly detection, localization, and type identification tasks. Our evaluation shows that state-of-the-art open-source and proprietary VLMs struggle on these specialized domains, and that existing visual prompting methods fail to reliably transfer to industrial anomaly analysis. To address this limitation, we propose ViPADIA, a lightweight visual prompting framework that uses inexpensive adversarially trained vision models to extract contextual hints in the form of visual and textual cues from industrial images. Across WTQA and CCQA, ViPADIA consistently outperforms prior visual prompting approaches and substantially improves VLM performance by up to 16.7 points in balanced accuracy and 10.9 points in macro F1. Moreover, by enabling low-cost VLMs to approach or even surpass the performance of much larger open-source and proprietary models, ViPADIA can substantially reduce inference costs while providing an efficient and practical solution for resource-constrained deployment in industrial inspection settings.