paper-with-me

홈 › Papers

Estimating the Number of Clusters via Normalized Cluster Instability

2016-08-26 · Jonas M. B. Haslbeck, Dirk U. Wulff

We improve current instability-based methods for the selection of the number of clusters $k$ in cluster analysis by developing a normalized cluster instability measure that corrects for the distribution of cluster sizes, a previously unaccounted driver of cluster instability. We show that our normalized instability measure outperforms current instability-based measures across the whole sequence of possible $k$ and especially overcomes limitations in the context of large $k$. We also compare, for the first time, model-based and model-free approaches to determine cluster-instability and find their performance to be comparable. We make our method available in the R-package \verb+cstab+.

📄 PDF Abstract BibTeX arXiv:1608.07494

Code (2)

cran/cstab
jmbh/cstab

Similar Papers 제목 키워드 기반

Bi-cross validation for estimating spectral clustering hyper parameters

2019-08-10 · Sioan Zohar, Chun-Hong Yoon

One challenge impeding the analysis of terabyte scale x-ray scattering data from the Linac Coherent Light Source LCLS, is determining the number of clusters required for the execution of traditional clustering algorithms…

Clustering

Distribution free optimality intervals for clustering

2021-07-30 · Marina Meilă, Hanyu Zhang

We address the problem of validating the ouput of clustering algorithms. Given data $\mathcal{D}$ and a partition $\mathcal{C}$ of these data into $K$ clusters, when can we say that the clusters obtained are correct or m…

Clustering

Penalized k-means algorithms for finding the correct number of clusters in a dataset

2019-11-15 · Behzad Kamgar-Parsi, Behrooz Kamgar-Parsi

In many applications we want to find the number of clusters in a dataset. A common approach is to use the penalized k-means algorithm with an additive penalty term linear in the number of clusters. An open problem is est…

Estimating the Mean Number of K-Means Clusters to Form

2015-03-07 · Robert A. Murphy

Utilizing the sample size of a dataset, the random cluster model is employed in order to derive an estimate of the mean number of K-Means clusters to form during classification of a dataset.

FormGeneral Classification

Approximating the clusters' prior distribution in Bayesian nonparametric models

2020-11-23 · pproximateinference AABI Symposium 2021 1 · Daria Bystrova, Julyan Arbel, Guillaume Kon Kam King, François Deslandes

In Bayesian nonparametrics, knowledge of the prior distribution induced on the number of clusters is key for prior specification and calibration. However, evaluating this prior is infamously difficult even for moderate s…