paper-with-me

홈 › Papers

Variable subset selection via GA and information complexity in mixtures of Poisson and negative binomial regression models

2015-05-20 · T. J. Massaro, H. Bozdogan

Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of regression models involving counts that come from a Poisson distribution, as well as a negative binomial distribution when the counts are significantly overdispersed. Within our proposed modeling framework, we carry out optimal component selection using the information criteria scores AIC, BIC, CAIC, and ICOMP. We demonstrate applications of our approach on simulated data, as well as on a real data set of HIV cases in Tennessee counties from the year 2010. Finally, using a genetic algorithm within our framework, we perform variable subset selection to determine the covariates that are most responsible for categorizing Tennessee counties. This leads to some interesting insights into the traits of counties that have high HIV counts.

📄 PDF Abstract BibTeX arXiv:1505.05229

Code (0)

등록된 구현이 없습니다.

Tasks

Epidemiologyregression

Similar Papers 제목 키워드 기반

Minimax Theory for High-dimensional Gaussian Mixtures with Sparse Mean Separation

2013-06-09 · NeurIPS 2013 12 · Martin Azizyan, Aarti Singh, Larry Wasserman

While several papers have investigated computationally and statistically efficient methods for learning Gaussian mixtures, precise minimax bounds for their statistical performance as well as fundamental limits in high-di…

Clusteringfeature selectionVocal Bursts Intensity Prediction

Subset selection in sparse matrices

2018-10-05 · Alberto Del Pia, Santanu S. Dey, Robert Weismantel

In subset selection we search for the best linear predictor that involves a small subset of variables. From a computational complexity viewpoint, subset selection is NP-hard and few classes are known to be solvable in po…

Higher Order Mutual Information Approximation for Feature Selection

2016-12-02 · Jilin Wu, Soumyajit Gupta, Chandrajit Bajaj

Feature selection is a process of choosing a subset of relevant features so that the quality of prediction models can be improved. An extensive body of work exists on information-theoretic feature selection, based on max…

feature selection

Feature Selection Facilitates Learning Mixtures of Discrete Product Distributions

2017-11-25 · Vincent Zhao, Steven W. Zucker

Feature selection can facilitate the learning of mixtures of discrete random variables as they arise, e.g. in crowdsourcing tasks. Intuitively, not all workers are equally reliable but, if the less reliable ones could be…

feature selection

Grouped Variable Selection for Generalized Eigenvalue Problems

2021-05-28 · Jonathan Dan, Simon Geirnaert, Alexander Bertrand

Many problems require the selection of a subset of variables from a full set of optimization variables. The computational complexity of an exhaustive search over all possible subsets of variables is, however, prohibitive…

Variable Selection