paper-with-me

Papers

Gaussian Mixture Clustering Using Relative Tests of Fit

2019-10-07 · Purvasha Chakravarti, Sivaraman Balakrishnan, Larry Wasserman

We consider clustering based on significance tests for Gaussian Mixture Models (GMMs). Our starting point is the SigClust method developed by Liu et al. (2008), which introduces a test based on the k-means objective (with k = 2) to decide whether the data should be split into two clusters. When applied recursively, this test yields a method for hierarchical clustering that is equipped with a significance guarantee. We study the limiting distribution and power of this approach in some examples and show that there are large regions of the parameter space where the power is low. We then introduce a new test based on the idea of relative fit. Unlike prior work, we test for whether a mixture of Gaussians provides a better fit relative to a single Gaussian, without assuming that either model is correct. The proposed test has a simple critical value and provides provable error control. One version of our test provides exact, finite sample control of the type I error. We show how our tests can be used for hierarchical clustering as well as in a sequential manner for model selection. We conclude with an extensive simulation study and a cluster analysis of a gene expression dataset.

📄 PDF Abstract BibTeX arXiv:1910.02566

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringModel Selection

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Confirmation of Binary Clustering in Gamma-Ray Bursts through an Integrated $p$-value from Multiple Nonparametric Tests of Hypotheses

2026-05-06 · Soumita Modak arxiv

The paper applies a new, nonparametric, interpoint distance-based measure to confirm the inherent groups prevailing in the brightest source of light in the universe: gamma-ray bursts. Our effective metric, in association…

The Infinite Mixture of Infinite Gaussian Mixtures

2014-12-01 · NeurIPS 2014 12 · Halid Z. Yerebakan, Bartek Rajwa, Murat Dundar

Dirichlet process mixture of Gaussians (DPMG) has been used in the literature for clustering and density estimation problems. However, many real-world data exhibit cluster distributions that cannot be captured by a singl…

ClusteringDensity Estimation

Spectral clustering in the Gaussian mixture block model

2023-04-29 · Shuangping Li, Tselil Schramm

Gaussian mixture block models are distributions over graphs that strive to model modern networks: to generate a graph from such a model, we associate each vertex $i$ with a latent feature vector $u_i \in \mathbb{R}^d$ sa…

Clusteringmodel

Clustering Semi-Random Mixtures of Gaussians

2017-11-23 · ICML 2018 7 · Pranjal Awasthi, Aravindan Vijayaraghavan

Gaussian mixture models (GMM) are the most widely used statistical model for the $k$-means clustering problem and form a popular framework for clustering in machine learning and data analysis. In this paper, we propose a…

Clustering

Vine copula mixture models and clustering for non-Gaussian data

2021-02-05 · Özge Sahin, Claudia Czado

The majority of finite mixture models suffer from not allowing asymmetric tail dependencies within components and not capturing non-elliptical clusters in clustering applications. Since vine copulas are very flexible in …

ClusteringModel Selectionparameter estimation