SF-DST: Adapting Vision-Language Models for Anomaly Detection via Asymmetric Modulation and A-LoRA
Abstract
Vision-Language Models (VLMs) show great potential for Industrial Anomaly Detection (IAD) but often overlook subtle high-frequency defects due to their semantic bias. Additionally, one-class adaptation risks rank collapse and catastrophic forgetting. To address these issues, we propose a unified framework termed SF-DST. First, we design Spatial-Frequency Dual-Stream Transformer, which explicitly captures lost details via a parallel discrete cosine transform branch. Crucially, it incorporates Asymmetric Cross-Modal Modulation (ACMM), a mechanism that leverages spatial semantics as a context gate to selectively retrieve and inject spectral features, effectively compensating for the texture blindness of VLMs. Second, Anomaly-Aware LoRA (A-LoRA) enforces Orthogonal Regularization on the adaptation subspace. This geometric constraint ensures subspace diversity to prevent mode collapse, while optimizing a hypersphere-based metric for precise anomaly discrimination. Extensive experiments on MVTec-AD, VisA and MMAD benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, successfully bridging the gap between large-model generalization and industrial-grade precision.