A new LDA formulation with covariates
The Latent Dirichlet Allocation (LDA) model is a popular method for creating mixed-membership clusters. Despite having been originally developed for text analysis, LDA has been used for a wide range of other applications. We propose a new formulation for the LDA model which incorporates covariates. In this model, a negative binomial regression is embedded within LDA, enabling straight-forward interpretation of the regression coefficients and the analysis of the quantity of cluster-specific elements in each sampling units (instead of the analysis being focused on modeling the proportion of each cluster, as in Structural Topic Models). We use slice sampling within a Gibbs sampling algorithm to estimate model parameters. We rely on simulations to show how our algorithm is able to successfully retrieve the true parameter values and the ability to make predictions for the abundance matrix using the information given by the covariates. The model is illustrated using real data sets from three different areas: text-mining of Coronavirus articles, analysis of grocery shopping baskets, and ecology of tree species on Barro Colorado Island (Panama). This model allows the identification of mixed-membership clusters in discrete data and provides inference on the relationship between covariates and the abundance of these clusters.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesregressionTopic ModelsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A convex formulation for high-dimensional sparse sliced inverse regression
Sliced inverse regression is a popular tool for sufficient dimension reduction, which replaces covariates with a minimal set of their linear combinations without loss of information on the conditional distribution of the…
Dimensionality ReductionregressionVariable SelectionVocal Bursts Intensity PredictionSurvival Permanental Processes for Survival Analysis with Time-Varying Covariates
Survival or time-to-event data with time-varying covariates are common in practice, and exploring the non-stationarity in covariates is essential to accurately analyzing the nonlinear dependence of time-to-event outcomes…
Partial Identification with Noisy Covariates: A Robust Optimization Approach
Causal inference from observational datasets often relies on measuring and adjusting for covariates. In practice, measurements of the covariates can often be noisy and/or biased, or only measurements of their proxies may…
Causal InferenceRanking and Selection with Covariates for Personalized Decision Making
We consider a problem of ranking and selection via simulation in the context of personalized decision making, where the best alternative is not universal but varies as a function of some observable covariates. The goal o…
Decision MakingExperimental DesignSpatially relaxed inference on high-dimensional linear models
We consider the inference problem for high-dimensional linear models, when covariates have an underlying spatial organization reflected in their correlation. A typical example of such a setting is high-resolution imaging…
ClusteringConstrained ClusteringVocal Bursts Intensity Prediction