What Actually Matters in Sensor Anomaly Detection? A Controlled Cross-Domain Study of Fusion, Temporal Context, and Metrics
Abstract
Background & motivation: Anomaly detection on multivariate sensor streams underpins applications from server monitoring to vehicle damage detection. A rich line of work has advanced detection on benchmarks such as the Server Machine Dataset (SMD) by proposing ever-new architectures, including stochastic recurrent models [2] and attention-based or reconstruction-forecasting hybrids, while a parallel line has shown these methods' headline metric, point-adjusted F1 (PA-F1), can be inflated even by random scoring [3]. Two questions fall between these lines. First, architecture comparisons change many factors at once (capacity, training scheme, and channel handling), so a basic design choice remains unisolated: should channels be encoded jointly (early fusion) or scored independently and pooled (late fusion)? Evidence from vehicle sensor data suggests fusion strategy matters [1], but a controlled test on a public benchmark is missing. Second, it is unknown whether the PA-F1 pathology merely inflates scores or actually reverses such design decisions. We address both with a deliberately minimal setup, using identical lightweight autoencoders (AEs) and varying one factor at a time, across two domains. Method: We compare five detectors on two open benchmarks from different domains: all 28 machines of the Server Machine Dataset (SMD; 38 server-telemetry channels each) and all 34 labeled runs of the SKAB industrial testbed (8 physical sensors: accelerometers, pressure, current, temperature, flow) [5]. Detectors: (i) a per-channel z-score baseline with max pooling; (ii) a dense AE encoding all channels of a single timestep jointly (early fusion); (iii) the same joint AE on 20-step windows, controlling for temporal context; and an ensemble of tiny per-channel windowed AEs (window 20) with (iv) mean or (v) max score pooling (late fusion). All models are lightweight (under 30 s training per unit on CPU). We report ROC-AUC and point-adjusted F1 [4], means ± std across units, with sign tests for pairwise comparisons. Results: Q1 (fusion vs. context): a naive comparison suggests a large late-fusion advantage on SMD: 0.914 ± 0.088 ROC-AUC (per-channel AEs, mean pooling) versus 0.809 ± 0.127 for a single-timestep joint AE. But this confounds fusion with temporal context: a joint AE on the same 20-step windows closes most of the gap (0.898 ± 0.099). The pattern replicates on SKAB, where windowing the joint AE adds 17 points (0.711 to 0.877). The residual late-fusion advantage, however, is domain-dependent: significant on SMD (+1.6 points, better on 21/28 machines, sign test p = 0.006) but absent on SKAB (0.873 vs. 0.877, 16/34, p = 0.70). Q2 (metrics): the ranking inverts under PA-F1 (on SMD: late fusion 0.804 vs. single-step joint AE 0.905), and on SKAB PA-F1 saturates entirely: the trivial z-score baseline attains 0.999 and every method exceeds 0.975, while ROC-AUC still separates them by 18 points. Across all runs the two metrics are essentially uncorrelated (r ≈ −0.1). The known pathology of point-adjusted evaluation is therefore not merely score inflation; it would lead a practitioner to select the wrong design [3]. Takeaway: Across two domains, the factors that robustly decide detection quality are temporal context and metric choice, not fusion strategy, whose residual effect is small and dataset-dependent. Uncontrolled comparisons and PA-F1-only evaluation can each reverse conclusions; sensor anomaly detectors should be compared with confound controls, across domains, and never on PA-F1 alone. References.: [1] Khan, S., et al. (2025). Robust anomaly detection through multi-modal autoencoder fusion for small vehicle damage detection. Machine Learning with Applications, Elsevier. [2] Su, Y., et al. (2019). Robust anomaly detection for multivariate time series through stochastic recurrent neural network. Proc. KDD '19. [3] Kim, S., et al. (2022). Towards a rigorous evaluation of time-series anomaly detection. Proc. AAAI 36(7). [4] Xu, H., et al. (2018). Unsupervised anomaly detection via variational auto-encoder for seasonal KPIs in web applications. Proc. WWW'18. [5] Katser, I. D., & Kozitsin, V. O. (2020). Skoltech Anomaly Benchmark (SKAB).