paper-with-me

홈 › Papers

An Aposteriorical Clusterability Criterion for $k$-Means++ and Simplicity of Clustering

2017-04-24 · Mieczysław A. Kłopotek

We define the notion of a well-clusterable data set combining the point of view of the objective of $k$-means clustering algorithm (minimising the centric spread of data elements) and common sense (clusters shall be separated by gaps). We identify conditions under which the optimum of $k$-means objective coincides with a clustering under which the data is separated by predefined gaps. We investigate two cases: when the whole clusters are separated by some gap and when only the cores of the clusters meet some separation condition. We overcome a major obstacle in using clusterability criteria due to the fact that known approaches to clusterability checking had the disadvantage that they are related to the optimal clustering which is NP hard to identify. Compared to other approaches to clusterability, the novelty consists in the possibility of an a posteriori (after running $k$-means) check if the data set is well-clusterable or not. As the $k$-means algorithm applied for this purpose has polynomial complexity so does therefore the appropriate check. Additionally, if $k$-means++ fails to identify a clustering that meets clusterability criteria, with high probability the data is not well-clusterable.

📄 PDF Abstract BibTeX arXiv:1704.07139

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringCommon Sense Reasoning

Similar Papers 제목 키워드 기반

Wide Gaps and Clustering Axioms

2023-08-07 · Mieczysław A. Kłopotek

The widely applied k-means algorithm produces clusterings that violate our expectations with respect to high/low similarity/density and is in conflict with Kleinberg's axiomatic system for distance based clustering algor…

Clustering

Clusterability-Based Assessment of Potentially Noisy Views for Multi-View Clustering

2026-04-20 · Mudi Jiang, Jiahui Zhou, Xinying Liu, Zengyou He 외 arxiv

In multi-view clustering, the quality of different views may vary substantially, and low-quality or degraded views can impair overall clustering performance. However, existing studies mainly address this issue within the…

Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review

2023-10-15 · Ravil Mussabayev, Rustam Mussabayev

This paper presents a comparative analysis of different optimization techniques for the K-means algorithm in the context of big data. K-means is a widely used clustering algorithm, but it can suffer from scalability issu…

Clustering

To Cluster, or Not to Cluster: An Analysis of Clusterability Methods

2018-08-24 · A. Adolfsson, M. Ackerman, N. C. Brownstein

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study o…

Clustering

Clusterability test for categorical data

2023-07-14 · Lianyu Hu, Junjie Dong, Mudi Jiang, Yan Liu 외

The objective of clusterability evaluation is to check whether a clustering structure exists within the data set. As a crucial yet often-overlooked issue in cluster analysis, it is essential to conduct such a test before…

AttributeClusteringvalid