An Efficient Cross-modal Feature Reconstruction Model for Multimodal Multi-class Anomaly Detection
Abstract
Industrial anomaly detection is evolving from category-specific detectors toward unified models capable of handling multiple categories simultaneously. However, in multi-class setting, greater inter-category diversity not only compels enhanced reconstruction capacity to model normal patterns but also exacerbates the “identical shortcut” problem, wherein anomalies are likewise well reconstructed. Moreover, existing unified frameworks are typically confined to single-modal (RGB) inputs and lack multimodal capability. To address these, we propose an Efficient Cross-modal Feature Reconstruction (ECFR) model for unified multi-class anomaly detection, which harnesses the inherent difficulty of cross-modal reconstruction to alleviate the shortcut issue and amplify anomaly discrimination. The framework is built on two core modules: an Adaptive Feature Interaction and Recalibration (AFIR) module and a Hybrid Attention Convolution (HAC) module. Through feature interaction, fusion, and reconstruction, our model achieves two key outcomes: learning normal patterns of multi-class objects while forcing reconstruction failures for anomalous inputs. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets validate our approach, which achieves state-of-the-art performance in multimodal multi-class anomaly detection while requiring significantly fewer model parameters (44.63 M) and lower computational complexity (15.77 GFLOPs) compared to previous multimodal anomaly detection methods.