paper-with-me

홈 › Papers

Non-parametric Power-law Data Clustering

2013-06-13 · Xuhui Fan, Yiling Zeng, Longbing Cao

It has always been a great challenge for clustering algorithms to automatically determine the cluster numbers according to the distribution of datasets. Several approaches have been proposed to address this issue, including the recent promising work which incorporate Bayesian Nonparametrics into the $k$-means clustering procedure. This approach shows simplicity in implementation and solidity in theory, while it also provides a feasible way to inference in large scale datasets. However, several problems remains unsolved in this pioneering work, including the power-law data applicability, mechanism to merge centers to avoid the over-fitting problem, clustering order problem, e.t.c.. To address these issues, the Pitman-Yor Process based k-means (namely \emph{pyp-means}) is proposed in this paper. Taking advantage of the Pitman-Yor Process, \emph{pyp-means} treats clusters differently by dynamically and adaptively changing the threshold to guarantee the generation of power-law clustering results. Also, one center agglomeration procedure is integrated into the implementation to be able to merge small but close clusters and then adaptively determine the cluster number. With more discussion on the clustering order, the convergence proof, complexity analysis and extension to spectral clustering, our approach is compared with traditional clustering algorithm and variational inference methods. The advantages and properties of pyp-means are validated by experiments on both synthetic datasets and real world datasets.

📄 PDF Abstract BibTeX arXiv:1306.3003

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringVariational Inference

Similar Papers 제목 키워드 기반

LazyPPL: laziness and types in non-parametric probabilistic programs

2021-10-08 · NeurIPS Workshop AIPLANS 2021 12 · Hugo Paquet, Sam Staton

We introduce LazyPPL, a prototype probabilistic programming library for Haskell. The library emphasises the clarifying power of types, and the connection between non-parametric, stochastic processes and lazy (call by nee…

ClusteringGaussian ProcessesPoint ProcessesProbabilistic Programming

A Bi-clustering Framework for Consensus Problems

2014-04-30 · Mariano Tepper, Guillermo Sapiro

We consider grouping as a general characterization for problems such as clustering, community detection in networks, and multiple parametric model estimation. We are interested in merging solutions from different groupin…

ClusteringCommunity Detection

Model-Based Clustering of Nonparametric Weighted Networks with Application to Water Pollution Analysis

2017-12-21 · Amal Agarwal, Lingzhou Xue

Water pollution is a major global environmental problem, and it poses a great environmental risk to public health and biological diversity. This work is motivated by assessing the potential environmental threat of coal m…

ClusteringDiversity

On a Theory of Nonparametric Pairwise Similarity for Clustering: Connecting Clustering to Classification

2014-12-01 · NeurIPS 2014 12 · Yingzhen Yang, Feng Liang, Shuicheng Yan, Zhangyang Wang 외

Pairwise clustering methods partition the data space into clusters by the pairwise similarity between data points. The success of pairwise clustering largely depends on the pairwise similarity function defined over the d…

ClusteringDensity EstimationGeneral ClassificationMulti-class Classification

Identifiability of Nonparametric Mixture Models and Bayes Optimal Clustering

2018-02-12 · Bryon Aragam, Chen Dan, Eric P. Xing, Pradeep Ravikumar

Motivated by problems in data clustering, we establish general conditions under which families of nonparametric mixture models are identifiable, by introducing a novel framework involving clustering overfitted \emph{para…

ClusteringNonparametric Clustering