Simultaneous Clustering and Model Selection for Multinomial Distribution: A Comparative Study
In this paper, we study different discrete data clustering methods, which use the Model-Based Clustering (MBC) framework with the Multinomial distribution. Our study comprises several relevant issues, such as initialization, model estimation and model selection. Additionally, we propose a novel MBC method by efficiently combining the partitional and hierarchical clustering techniques. We conduct experiments on both synthetic and real data and evaluate the methods using accuracy, stability and computation time. Our study identifies appropriate strategies to be used for discrete data analysis with the MBC methods. Moreover, our proposed method is very competitive w.r.t. clustering accuracy and better w.r.t. stability and computation time.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringModel SelectionSimilar Papers 제목 키워드 기반
A model selection approach for clustering a multinomial sequence with non-negative factorization
We consider a problem of clustering a sequence of multinomial observations by way of a model selection criterion. We propose a form of a penalty term for the model selection procedure. Our approach subsumes both the conv…
ClusteringModel SelectionMultiple co-clustering based on nonparametric mixture models with heterogeneous marginal distributions
We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is…
ClusteringVariational InferenceSimultaneous Dimensionality and Complexity Model Selection for Spectral Graph Clustering
Our problem of interest is to cluster vertices of a graph by identifying underlying community structure. Among various vertex clustering approaches, spectral clustering is one of the most popular methods because it is ea…
ClusteringGraph ClusteringModel SelectionSpectral Graph Clustering+1Automatic Response Category Combination in Multinomial Logistic Regression
We propose a penalized likelihood method that simultaneously fits the multinomial logistic regression model and combines subsets of the response categories. The penalty is non differentiable when pairs of columns in the …
Model SelectionregressionHierarchical mixtures of Unigram models for short text clustering: The role of Beta-Liouville priors
This paper presents a variant of the Multinomial mixture model tailored to the unsupervised classification of short text data. While the Multinomial probability vector is traditionally assigned a Dirichlet prior distribu…
Short Text ClusteringText ClusteringVariational Inference