Internal Evaluation of Unsupervised Anomaly Detection Algorithms with Explanation
Abstract
Although numerous algorithms have been proposed for unsupervised anomaly detection, there is currently no widely acknowledged internal evaluation method tailored specifically to this task. Existing internal evaluation methods typically assess score distributions alone or depend on auxiliary classifiers, limiting their reliability and interpretability. To address these problems, we propose SSSD (Similar Scores, Similar Data), an evaluation framework inspired by the smoothness assumption in machine learning—that similar inputs should yield similar outputs. We adapt this principle to anomaly detection by requiring that locally similar data points be assigned similar anomaly scores. Instead of testing this assumption directly, SSSD identifies violations of this principle by comparing the similarity of data with similar anomaly scores, interpreting a significant difference as a higher probability of erroneous detection. This evaluation of similarity is measured through nearby-score and nearby-distance, which is calculated based on the distances between neighboring data points and their corresponding anomaly scores. SSSD is highly interpretable, providing valuable insights for identifying mislabeled data. Experimental results show that our proposed method is both effective and robust for evaluating the performance of unsupervised anomaly detection algorithms across artificial and real-world datasets.