paper-with-me

홈 › Papers

A sparse negative binomial mixture model for clustering RNA-seq count data

2019-12-05 · Tanbin Rahman, Yujia Li, Tianzhou Ma, Lu Tang, George Tseng

Clustering with variable selection is a challenging yet critical task for modern small-n-large-p data. Existing methods based on sparse Gaussian mixture models or sparse K-means provide solutions to continuous data. With the prevalence of RNA-seq technology and lack of count data modeling for clustering, the current practice is to normalize count expression data into continuous measures and apply existing models with Gaussian assumption. In this paper, we develop a negative binomial mixture model with lasso or fused lasso gene regularization to cluster samples (small n) with high-dimensional gene features (large p). EM algorithm and Bayesian information criterion are used for inference and determining tuning parameters. The method is compared with existing methods using extensive simulations and two real transcriptomic applications in rat brain and breast cancer studies. The result shows superior performance of the proposed count data model in clustering accuracy, feature selection and biological interpretation in pathways.

📄 PDF Abstract BibTeX arXiv:1912.02399

Code (1)

YujiaLi1994/snbClust 공식 구현

Tasks

Clusteringfeature selectionVariable Selection

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

Generalized Negative Binomial Processes and the Representation of Cluster Structures

2013-10-07 · Mingyuan Zhou

The paper introduces the concept of a cluster structure to define a joint distribution of the sample size and its exchangeable random partitions. The cluster structure allows the probability distribution of the random pa…

Clustering

Fully Bayesian inference for neural models with negative-binomial spiking

2012-12-01 · NeurIPS 2012 12 · James Scott, Jonathan W. Pillow

Characterizing the information carried by neural populations in the brain requires accurate statistical models of neural spike responses. The negative-binomial distribution provides a convenient model for over-dispersed…

Bayesian InferenceData Augmentationregression

Optimal Clustering of Discrete Mixtures: Binomial, Poisson, Block Models, and Multi-layer Networks

2023-11-27 · Zhongyuan Lyu, Ting Li, Dong Xia

In this paper, we first study the fundamental limit of clustering networks when a multi-layer network is present. Under the mixture multi-layer stochastic block model (MMSBM), we show that the minimax optimal network clu…

ClusteringCommunity DetectionStochastic Block Model

Variable subset selection via GA and information complexity in mixtures of Poisson and negative binomial regression models

2015-05-20 · T. J. Massaro, H. Bozdogan

Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in wh…

Epidemiologyregression

Universal Lower Bounds and Optimal Rates: Achieving Minimax Clustering Error in Sub-Exponential Mixture Models

2024-02-23 · Maximilien Dreveton, Alperen Gözeten, Matthias Grossglauser, Patrick Thiran

Clustering is a pivotal challenge in unsupervised machine learning and is often investigated through the lens of mixture models. The optimal error rate for recovering cluster labels in Gaussian and sub-Gaussian mixture m…

Clustering