paper-with-me

Papers

Bayesian Cluster Enumeration Criterion for Unsupervised Learning

2017-10-22 · Freweyni K. Teklehaymanot, Michael Muma, Abdelhak M. Zoubir

We derive a new Bayesian Information Criterion (BIC) by formulating the problem of estimating the number of clusters in an observed data set as maximization of the posterior probability of the candidate models. Given that some mild assumptions are satisfied, we provide a general BIC expression for a broad class of data distributions. This serves as a starting point when deriving the BIC for specific distributions. Along this line, we provide a closed-form BIC expression for multivariate Gaussian distributed variables. We show that incorporating the data structure of the clustering problem into the derivation of the BIC results in an expression whose penalty term is different from that of the original BIC. We propose a two-step cluster enumeration algorithm. First, a model-based unsupervised learning algorithm partitions the data according to a given set of candidate models. Subsequently, the number of clusters is determined as the one associated with the model for which the proposed BIC is maximal. The performance of the proposed two-step algorithm is tested using synthetic and real data sets.

📄 PDF Abstract BibTeX arXiv:1710.07954

Code (1)

FreTekle/Bayesian-Cluster-Enumeration 공식 구현

Tasks

Clustering

Similar Papers 제목 키워드 기반

Robust Bayesian Cluster Enumeration Based on the $t$ Distribution

2018-11-29 · Freweyni K. Teklehaymanot, Michael Muma, Abdelhak M. Zoubir

A major challenge in cluster analysis is that the number of data clusters is mostly unknown and it must be estimated prior to clustering the observed data. In real-world applications, the observed data is often subject t…

Clustering

Robust M-Estimation Based Bayesian Cluster Enumeration for Real Elliptically Symmetric Distributions

2020-05-04 · Christian A. Schroth, Michael Muma

Robustly determining the optimal number of clusters in a data set is an essential factor in a wide range of applications. Cluster enumeration becomes challenging when the true underlying structure in the observed data is…

Person Identification

Approach of variable clustering and compression for learning large Bayesian networks

2022-08-29 · Anna V. Bubnova

This paper describes a new approach for learning structures of large Bayesian networks based on blocks resulting from feature space clustering. This clustering is obtained using normalized mutual information. And the sub…

Clustering

Un modèle Bayésien de co-clustering de données mixtes

2019-02-06 · Aichetou Bouchareb, Marc Boullé, Fabrice Rossi, Fabrice Clérot

We propose a MAP Bayesian approach to perform and evaluate a co-clustering of mixed-type data tables. The proposed model infers an optimal segmentation of all variables then performs a co-clustering by minimizing a Bayes…

ClusteringModel Selection

Interpretable Clustering with the Distinguishability Criterion

2024-04-24 · Ali Turfah, Xiaoquan Wen

Cluster analysis is a popular unsupervised learning tool used in many disciplines to identify heterogeneous sub-populations within a sample. However, validating cluster analysis results and determining the number of clus…

Clustering