paper-with-me

홈 › Papers

Asymptotics for The $k$-means

2022-11-18 · Tonglin Zhang

The $k$-means is one of the most important unsupervised learning techniques in statistics and computer science. The goal is to partition a data set into many clusters, such that observations within clusters are the most homogeneous and observations between clusters are the most heterogeneous. Although it is well known, the investigation of the asymptotic properties is far behind, leading to difficulties in developing more precise $k$-means methods in practice. To address this issue, a new concept called clustering consistency is proposed. Fundamentally, the proposed clustering consistency is more appropriate than the previous criterion consistency for the clustering methods. Using this concept, a new $k$-means method is proposed. It is found that the proposed $k$-means method has lower clustering error rates and is more robust to small clusters and outliers than existing $k$-means methods. When $k$ is unknown, using the Gap statistics, the proposed method can also identify the number of clusters. This is rarely achieved by existing $k$-means methods adopted by many software packages.

📄 PDF Abstract BibTeX arXiv:2211.10015

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

JUMP-Means: Small-Variance Asymptotics for Markov Jump Processes

2015-03-01 · Jonathan H. Huggins, Karthik Narasimhan, Ardavan Saeedi, Vikash K. Mansinghka

Markov jump processes (MJPs) are used to model a wide range of phenomena from disease progression to RNA path folding. However, maximum likelihood estimation of parametric models leads to degenerate trajectories and infe…

Small-Variance Asymptotics for Exponential Family Dirichlet Process Mixture Models

2012-12-01 · NeurIPS 2012 12 · Ke Jiang, Brian Kulis, Michael. I. Jordan

Links between probabilistic and non-probabilistic learning algorithms can arise by performing small-variance asymptotics, i.e., letting the variance of particular distributions in a graphical model go to zero. For instan…

Clustering

Small-Variance Asymptotics for Hidden Markov Models

2013-12-01 · NeurIPS 2013 12 · Anirban Roychowdhury, Ke Jiang, Brian Kulis

Small-variance asymptotics provide an emerging technique for obtaining scalable combinatorial algorithms from rich probabilistic models. We present a small-variance asymptotic analysis of the Hidden Markov Model and its …

Learning Gaussian Mixtures with Generalised Linear Models: Precise Asymptotics in High-dimensions

2021-06-07 · Bruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco 외

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of $K$ Gaussian…

ClassificationMulti-class ClassificationVocal Bursts Intensity Prediction

Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensions

2021-12-01 · NeurIPS 2021 12 · Bruno Loureiro, Gabriele Sicuro, Cedric Gerbelot, Alessandro Pacco 외

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of $K$ Gaussian…

ClassificationMulti-class Classification