paper-with-me

Papers

Homogeneity of Cluster Ensembles

2016-02-08 · Brijnesh J. Jain

The expectation and the mean of partitions generated by a cluster ensemble are not unique in general. This issue poses challenges in statistical inference and cluster stability. In this contribution, we state sufficient conditions for uniqueness of expectation and mean. The proposed conditions show that a unique mean is neither exceptional nor generic. To cope with this issue, we introduce homogeneity as a measure of how likely is a unique mean for a sample of partitions. We show that homogeneity is related to cluster stability. This result points to a possible conflict between cluster stability and diversity in consensus clustering. To assess homogeneity in a practical setting, we propose an efficient way to compute a lower bound of homogeneity. Empirical results using the k-means algorithm suggest that uniqueness of the mean partition is not exceptional for real-world data. Moreover, for samples of high homogeneity, uniqueness can be enforced by increasing the number of data points or by removing outlier partitions. In a broader context, this contribution can be placed as a further step towards a statistical theory of partitions.

📄 PDF Abstract BibTeX arXiv:1602.02543

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDiversity

Similar Papers 제목 키워드 기반

External Clustering Validation by the Homogeneity-Parsimony Trade-off

2026-07-22 · Andreas Tiffeau-Mayer arxiv

Scalar metrics are often used to evaluate clusterings against known classes, but they can obscure a fundamental trade-off: clusterings should be informative about class labels while avoiding unnecessary fragmentation. He…

Adaptive Nonparametric Clustering

2017-09-26 · Kirill Efimov, Larisa Adamyan, Vladimir Spokoiny

This paper presents a new approach to non-parametric cluster analysis called Adaptive Weights Clustering (AWC). The idea is to identify the clustering structure by checking at different points and for different scales on…

ClusteringNonparametric Clustering

Clustering Mixed Datasets Using Homogeneity Analysis with Applications to Big Data

2016-08-17 · Rajiv Sambasivan, Sourish Das

Datasets with a mixture of numerical and categorical attributes are routinely encountered in many application domains. In this work we examine an approach to clustering such datasets using homogeneity analysis. Homogenei…

Clustering

A New Homogeneity Inter-Clusters Measure in SemiSupervised Clustering

2013-04-13 · Badreddine Meftahi, Ourida Ben Boubaker Saidi

Many studies in data mining have proposed a new learning called semi-Supervised. Such type of learning combines unlabeled and labeled data which are hard to obtain. However, in unsupervised methods, the only unlabeled da…

Clustering

Diversifying Deep Ensembles: A Saliency Map Approach for Enhanced OOD Detection, Calibration, and Accuracy

2023-05-19 · Stanislav Dereka, Ivan Karpukhin, Maksim Zhdanov, Sergey Kolesnikov

Deep ensembles are capable of achieving state-of-the-art results in classification and out-of-distribution (OOD) detection. However, their effectiveness is limited due to the homogeneity of learned patterns within ensemb…

ClassificationDiversityOut of Distribution (OOD) Detection