AirIAD: Agentic Iterative Reasoning for Industrial Anomaly Detection
Abstract
Multimodal large language models offer a promising foundation for industrial anomaly detection (IAD), yet they remain unreliable for fine-grained diagnosis, where subtle defects require tight alignment between localized visual evidence and domain-specific semantics. Existing single-pass frameworks produce predictions without revisiting or validating intermediate evidence, which leads to spatio-semantic misalignment and error accumulation. More importantly, they lack mechanisms to actively control how additional evidence is acquired and integrated during reasoning, limiting their ability to resolve ambiguous cases. To address these limitations, we propose AirIAD, an agentic iterative reasoning framework that unifies progressive refinement with active information acquisition. AirIAD equips the model with two tools, Perceptive Zoomer and Knowledge Retriever, which enable it to iteratively gather complementary visual and semantic evidence. Through this process, the model performs explicit cross-verification between what is observed and what is known, refining intermediate hypotheses until convergence on a reliable diagnosis. This capability is enabled by a two-stage training pipeline. A supervised fine-tuning stage introduces a Spatio-Semantic Cross-Verification Chain-of-Thought, which structures coordinated perception and knowledge grounding. This is followed by a Spatio-Semantic Group-in-Group Policy Optimization stage, which provides dense step-wise supervision to encourage consistent self-refinement and robust spatio-semantic alignment. Extensive experiments on the MMAD benchmark show that AirIAD, built on Qwen3-VL-Instruct-4B, achieves state-of-the-art performance in anomaly detection and fine-grained diagnosis.