paper-with-me

홈 › Papers

CNMBI: Determining the Number of Clusters Using Center Pairwise Matching and Boundary Filtering

2026-03-23 · Ruilin Zhang, Haiyang Zheng, Hongpeng Wang arxiv

One of the main challenges in data mining is choosing the optimal number of clusters without prior information. Notably, existing methods are usually in the philosophy of cluster validation and hence have underlying assumptions on data distribution, which prevents their application to complex data such as large-scale images and high-dimensional data from the real world. In this regard, we propose an approach named CNMBI. Leveraging the distribution information inherent in the data space, we map the target task as a dynamic comparison process between cluster centers regarding positional behavior, without relying on the complete clustering results and designing the complex validity index as before. Bipartite graph theory is then employed to efficiently model this process. Additionally, we find that different samples have different confidence levels and thereby actively remove low-confidence ones, which is, for the first time to our knowledge, considered in cluster number determination. CNMBI is robust and allows for more flexibility in the dimension and shape of the target data (e.g., CIFAR-10 and STL-10). Extensive comparison studies with state-of-the-art competitors on various challenging datasets demonstrate the superiority of our method.

📄 PDF Abstract BibTeX arXiv:2603.26744

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Estimating the Optimal Number of Clusters in Categorical Data Clustering by Silhouette Coefficient

2025-01-26 · Duy-Tai Dinh, Tsutomu Fujinami, Van-Nam Huynh

The problem of estimating the number of clusters (say k) is one of the major challenges for the partitional clustering. This paper proposes an algorithm named k-SCC to estimate the optimal k in categorical data clusterin…

Categorical data clusteringClusteringDensity Estimation

Neural network-based clustering using pairwise constraints

2015-11-19 · Yen-Chang Hsu, Zsolt Kira

This paper presents a neural network-based end-to-end clustering framework. We design a novel strategy to utilize the contrastive criteria for pushing data-forming clusters directly from raw data, in addition to learning…

Clustering

Cluster validity index based on Jeffrey divergence

2018-12-20 · Ahmed Ben Said, Rachid Hadjidj, Sebti Foufou

Cluster validity indexes are very important tools designed for two purposes: comparing the performance of clustering algorithms and determining the number of clusters that best fits the data. These indexes are in general…

Clustering

Convergence Analysis of Gradient EM for Multi-component Gaussian Mixture

2017-05-23 · Bowei Yan, Mingzhang Yin, Purnamrita Sarkar

In this paper, we study convergence properties of the gradient Expectation-Maximization algorithm \cite{lange1995gradient} for Gaussian Mixture Models for general number of clusters and mixing coefficients. We derive the…

Learning Theory

Convergence of Gradient EM on Multi-component Mixture of Gaussians

2017-12-01 · NeurIPS 2017 12 · Bowei Yan, Mingzhang Yin, Purnamrita Sarkar

In this paper, we study convergence properties of the gradient variant of Expectation-Maximization algorithm~\cite{lange1995gradient} for Gaussian Mixture Models for arbitrary number of clusters and mixing coefficients. …

Learning Theory