paper-with-me

홈 › Papers

Clustering by the Probability Distributions from Extreme Value Theory

2022-02-20 · Sixiao Zheng, Ke Fan, Yanxi Hou, Jianfeng Feng, Yanwei Fu

Clustering is an essential task to unsupervised learning. It tries to automatically separate instances into coherent subsets. As one of the most well-known clustering algorithms, k-means assigns sample points at the boundary to a unique cluster, while it does not utilize the information of sample distribution or density. Comparably, it would potentially be more beneficial to consider the probability of each sample in a possible cluster. To this end, this paper generalizes k-means to model the distribution of clusters. Our novel clustering algorithm thus models the distributions of distances to centroids over a threshold by Generalized Pareto Distribution (GPD) in Extreme Value Theory (EVT). Notably, we propose the concept of centroid margin distance, use GPD to establish a probability model for each cluster, and perform a clustering algorithm based on the covering probability function derived from GPD. Such a GPD k-means thus enables the clustering algorithm from the probabilistic perspective. Correspondingly, we also introduce a naive baseline, dubbed as Generalized Extreme Value (GEV) k-means. GEV fits the distribution of the block maxima. In contrast, the GPD fits the distribution of distance to the centroid exceeding a sufficiently large threshold, leading to a more stable performance of GPD k-means. Notably, GEV k-means can also estimate cluster structure and thus perform reasonably well over classical k-means. Thus, extensive experiments on synthetic datasets and real datasets demonstrate that GPD k-means outperforms competitors. The github codes are released in https://github.com/sixiaozheng/EVT-K-means.

📄 PDF Abstract BibTeX arXiv:2202.09784

Code (1)

sixiaozheng/evt-k-means 공식 구현 pytorch

Tasks

Clustering

Similar Papers 제목 키워드 기반

Extreme Value k-means Clustering

2019-09-25 · Sixiao Zheng, Yanxi Hou, Yanwei Fu, Jianfeng Feng

Clustering is the central task in unsupervised learning and data mining. k-means is one of the most widely used clustering algorithms. Unfortunately, it is generally non-trivial to extend k-means to cluster data points b…

ClusteringComputational Efficiency

A Multivariate Extreme Value Theory Approach to Anomaly Clustering and Visualization

2019-07-17 · Maël Chiapino, Stéphan Clémençon, Vincent Feuillard, Anne Sabourin

In a wide variety of situations, anomalies in the behaviour of a complex system, whose health is monitored through the observation of a random vector X = (X1,. .. , X d) valued in R d , correspond to the simultaneous occ…

ClusteringGraph Mining

Estimation of tail risk measures in finance: Approaches to extreme value mixture modeling

2024-06-01 · Yujuan Qiu

This thesis evaluates most of the extreme mixture models and methods that have appended in the literature and implements them in the context of finance and insurance. The paper also reviews and studies extreme value theo…

Density EstimationTime Series

Evaluating probabilistic forecasts of extremes using continuous ranked probability score distributions

2019-05-10 · Maxime Taillardat, Anne-Laure Fougères, Philippe Naveau, Raphaël de Fondeville

Verifying probabilistic forecasts for extreme events is a highly active research area because popular media and public opinions are naturally focused on extreme events, and biased conclusions are readily made. In this co…

ExGAN: Adversarial Generation of Extreme Samples

2020-09-17 · Siddharth Bhatia, Arjit Jain, Bryan Hooi

Mitigating the risk arising from extreme events is a fundamental goal with many applications, such as the modelling of natural disasters, financial crashes, epidemics, and many others. To manage this risk, a vital step i…

Extreme Sample Generation