paper-with-me

홈 › Papers

CFOF: A Concentration Free Measure for Anomaly Detection

2019-01-14 · Fabrizio Angiulli

We present a novel notion of outlier, called the Concentration Free Outlier Factor, or CFOF. As a main contribution, we formalize the notion of concentration of outlier scores and theoretically prove that CFOF does not concentrate in the Euclidean space for any arbitrary large dimensionality. To the best of our knowledge, there are no other proposals of data analysis measures related to the Euclidean distance for which it has been provided theoretical evidence that they are immune to the concentration effect. We determine the closed form of the distribution of CFOF scores in arbitrarily large dimensionalities and show that the CFOF score of a point depends on its squared norm standard score and on the kurtosis of the data distribution, thus providing a clear and statistically founded characterization of this notion. Moreover, we leverage this closed form to provide evidence that the definition does not suffer of the hubness problem affecting other measures. We prove that the number of CFOF outliers coming from each cluster is proportional to cluster size and kurtosis, a property that we call semi-locality. We determine that semi-locality characterizes existing reverse nearest neighbor-based outlier definitions, thus clarifying the exact nature of their observed local behavior. We also formally prove that classical distance-based and density-based outliers concentrate both for bounded and unbounded sample sizes and for fixed and variable values of the neighborhood parameter. We introduce the fast-CFOF algorithm for detecting outliers in large high-dimensional dataset. The algorithm has linear cost, supports multi-resolution analysis, and is embarrassingly parallel. Experiments highlight that the technique is able to efficiently process huge datasets and to deal even with large values of the neighborhood parameter, to avoid concentration, and to obtain excellent accuracy.

📄 PDF Abstract BibTeX arXiv:1901.04992

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detection

Similar Papers 제목 키워드 기반

Batch Distillation Data for Developing Machine Learning Anomaly Detection Methods

2025-10-20 · Justus Arweiler, Indra Jungjohann, Aparna Muraleedharan, Heike Leitte 외 arxiv

Machine learning (ML) holds great potential to advance anomaly detection (AD) in chemical processes. However, the development of ML-based methods is hindered by the lack of openly available experimental data. To address …

Anomaly Detection

A Deep Learning Approach to Anomaly Sequence Detection for High-Resolution Monitoring of Power Systems

2020-12-09 · Kursat Rasim Mestav, Xinyi Wang, Lang Tong

A deep learning approach is proposed to detect data and system anomalies using high-resolution continuous point-on-wave (CPOW) or phasor measurements. Both the anomaly and anomaly-free measurement models are assumed to h…

Anomaly DetectionGenerative Adversarial Network

Early Cancer Detection in Blood Vessels Using Mobile Nanosensors

2018-05-22

In this paper, we propose using mobile nanosensors (MNSs) for early stage anomaly detection. For concreteness, we focus on the detection of cancer cells located in a particular region of a blood vessel. These cancer cell…

Anomaly Detection

Robust Spectral Anomaly Detection in EELS Spectral Images via Three Dimensional Convolutional Variational Autoencoders

2024-12-16 · Seyfal Sultanov, James P Buban, Robert F Klie

We introduce a Three-Dimensional Convolutional Variational Autoencoder (3D-CVAE) for automated anomaly detection in Electron Energy Loss Spectroscopy Spectrum Imaging (EELS-SI) data. Our approach leverages the full three…

Anomaly Detection

Universal Data Anomaly Detection via Inverse Generative Adversary Network

2020-01-23 · Kursat Rasim Mestav, Lang Tong

The problem of detecting data anomaly is considered. Under the null hypothesis that models anomaly-free data, measurements are assumed to be from an unknown distribution with some authenticated historical samples. Under …

Anomaly Detection