paper-with-me

Papers

Statistical Inference for Clustering-based Anomaly Detection

2025-04-25 · Nguyen Thi Minh Phu, Duong Tan Loc, Vo Nguyen Le Duy

Unsupervised anomaly detection (AD) is a fundamental problem in machine learning and statistics. A popular approach to unsupervised AD is clustering-based detection. However, this method lacks the ability to guarantee the reliability of the detected anomalies. In this paper, we propose SI-CLAD (Statistical Inference for CLustering-based Anomaly Detection), a novel statistical framework for testing the clustering-based AD results. The key strength of SI-CLAD lies in its ability to rigorously control the probability of falsely identifying anomalies, maintaining it below a pre-specified significance level $\alpha$ (e.g., $\alpha = 0.05$). By analyzing the selection mechanism inherent in clustering-based AD and leveraging the Selective Inference (SI) framework, we prove that false detection control is attainable. Moreover, we introduce a strategy to boost the true detection rate, enhancing the overall performance of SI-CLAD. Extensive experiments on synthetic and real-world datasets provide strong empirical support for our theoretical findings, showcasing the superior performance of the proposed method.

📄 PDF Abstract BibTeX arXiv:2504.18633

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly DetectionClusteringUnsupervised Anomaly Detection

Similar Papers 제목 키워드 기반

Statistically Significant $k$NNAD by Selective Inference

2025-02-18 · Mizuki Niihori, Teruyuki Katsuoka, Tomohiro Shiraishi, Shuichi Nishino 외

In this paper, we investigate the problem of unsupervised anomaly detection using the k-Nearest Neighbor method. The k-Nearest Neighbor Anomaly Detection (kNNAD) is a simple yet effective approach for identifying anomali…

Anomaly DetectionUnsupervised Anomaly Detection

Multi-level conformal clustering: A distribution-free technique for clustering and anomaly detection

2019-10-17 · Ilia Nouretdinov, James Gammerman, Matteo Fontana, Daljit Rehal

In this work we present a clustering technique called \textit{multi-level conformal clustering (MLCC)}. The technique is hierarchical in nature because it can be performed at multiple significance levels which yields gre…

Anomaly DetectionClusteringConformal Prediction

Network Anomaly Detection: A Survey and Comparative Analysis of Stochastic and Deterministic Methods

2013-09-19 · Jing Wang, Daniel Rossell, Christos G. Cassandras, Ioannis Ch. Paschalidis

We present five methods to the problem of network anomaly detection. These methods cover most of the common techniques in the anomaly detection field, including Statistical Hypothesis Tests (SHT), Support Vector Machines…

Anomaly DetectionClustering

Multivariate Time series Anomaly Detection:A Framework of Hidden Markov Models

2025-11-11 · Jinbo Li, Witold Pedrycz, Iqbal Jamal arxiv

In this study, we develop an approach to multivariate time series anomaly detection focused on the transformation of multivariate time series to univariate time series. Several transformation techniques involving Fuzzy C…

Time Series Anomaly Detection

Statistical Inference Using Mean Shift Denoising

2016-10-13 · Yunhua Xiang, Yen-Chi Chen

In this paper, we study how the mean shift algorithm can be used to denoise a dataset. We introduce a new framework to analyze the mean shift algorithm as a denoising approach by viewing the algorithm as an operator on a…

Anomaly DetectionClusteringDenoising