paper-with-me

Papers

Robust Bayesian Cluster Enumeration Based on the $t$ Distribution

2018-11-29 · Freweyni K. Teklehaymanot, Michael Muma, Abdelhak M. Zoubir

A major challenge in cluster analysis is that the number of data clusters is mostly unknown and it must be estimated prior to clustering the observed data. In real-world applications, the observed data is often subject to heavy tailed noise and outliers which obscure the true underlying structure of the data. Consequently, estimating the number of clusters becomes challenging. To this end, we derive a robust cluster enumeration criterion by formulating the problem of estimating the number of clusters as maximization of the posterior probability of multivariate $t_\nu$ distributed candidate models. We utilize Bayes' theorem and asymptotic approximations to come up with a robust criterion that possesses a closed-form expression. Further, we refine the derivation and provide a robust cluster enumeration criterion for data sets with finite sample size. The robust criteria require an estimate of cluster parameters for each candidate model as an input. Hence, we propose a two-step cluster enumeration algorithm that uses the expectation maximization algorithm to partition the data and estimate cluster parameters prior to the calculation of one of the robust criteria. The performance of the proposed algorithm is tested and compared to existing cluster enumeration methods using numerical and real data experiments.

📄 PDF Abstract BibTeX arXiv:1811.12337

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Robust M-Estimation Based Bayesian Cluster Enumeration for Real Elliptically Symmetric Distributions

2020-05-04 · Christian A. Schroth, Michael Muma

Robustly determining the optimal number of clusters in a data set is an essential factor in a wide range of applications. Cluster enumeration becomes challenging when the true underlying structure in the observed data is…

Person Identification

Bayesian Cluster Enumeration Criterion for Unsupervised Learning

2017-10-22 · Freweyni K. Teklehaymanot, Michael Muma, Abdelhak M. Zoubir

We derive a new Bayesian Information Criterion (BIC) by formulating the problem of estimating the number of clusters in an observed data set as maximization of the posterior probability of the candidate models. Given tha…

Clustering

Approach of variable clustering and compression for learning large Bayesian networks

2022-08-29 · Anna V. Bubnova

This paper describes a new approach for learning structures of large Bayesian networks based on blocks resulting from feature space clustering. This clustering is obtained using normalized mutual information. And the sub…

Clustering

Uniform random generation of large acyclic digraphs

2012-02-29 · Jack Kuipers, Giusi Moffa

Directed acyclic graphs are the basic representation of the structure underlying Bayesian networks, which represent multivariate probability distributions. In many practical applications, such as the reverse engineering …

Solution Enumeration by Optimality in Answer Set Programming

2021-08-07 · Jukka Pajunen, Tomi Janhunen

Given a combinatorial search problem, it may be highly useful to enumerate its (all) solutions besides just finding one solution, or showing that none exists. The same can be stated about optimal solutions if an objectiv…