FACETS: Cross-Granularity Vision--Language Modeling for 3D Anomaly Detection
Abstract
Three-dimensional (3D) anomaly detection underpins quality assurance in advanced manufacturing and engineering, capturing subtle defects and deviations that 2D inspection struggles to resolve due to occlusion, viewpoint, and appearance confounds. Existing 3D anomaly detection methods either rely on reconstruction errors or stored normal representations to detect anomalies. While both paradigms have driven substantial progress, they share a fundamental limitation: both operate entirely in the visual domain, missing the rich semantics encoded in natural language, which also typically constrains them to per-category models. We propose FACETS, among the first frameworks that leverage a 3D vision--language model for 3D anomaly detection, opening a new paradigm beyond reconstruction- and memory-bank-based methods. The key idea is to explicitly retain native point-level features that enable reasoning at both patch and point granularities, and further leverage language grounding and cross-granularity geometric modeling along two complementary axes, linguistic semantics and geometric saliency, so that coarse semantic cues and fine-grained geometric details jointly support anomaly detection. FACETS enables unified multi-category anomaly detection, avoiding the per-category models used by most prior methods. Notably, we provide a mathematical analysis of our loss function, offering valuable insights into the substantial improvement FACETS achieves in anomaly localization. Extensive experiments on popular benchmarks reveal that FACETS establishes new SOTA performance across datasets and metrics for 3D anomaly detection, substantially and consistently outperforming existing methods. Code is provided as supplementary material.