paper-with-me

Papers

DOLDA - a regularized supervised topic model for high-dimensional multi-class regression

2016-01-31 · Måns Magnusson, Leif Jonsson, Mattias Villani

Generating user interpretable multi-class predictions in data rich environments with many classes and explanatory covariates is a daunting task. We introduce Diagonal Orthant Latent Dirichlet Allocation (DOLDA), a supervised topic model for multi-class classification that can handle both many classes as well as many covariates. To handle many classes we use the recently proposed Diagonal Orthant (DO) probit model (Johndrow et al., 2013) together with an efficient Horseshoe prior for variable selection/shrinkage (Carvalho et al., 2010). We propose a computationally efficient parallel Gibbs sampler for the new model. An important advantage of DOLDA is that learned topics are directly connected to individual classes without the need for a reference class. We evaluate the model's predictive accuracy on two datasets and demonstrate DOLDA's advantage in interpreting the generated predictions.

📄 PDF Abstract BibTeX arXiv:1602.00260

Code (1)

lejon/DiagonalOrthantLDA 공식 구현

Tasks

General ClassificationMulti-class ClassificationregressionVariable Selection

Similar Papers 제목 키워드 기반

High-Dimensional Tail Index Regression: with An Application to Text Analyses of Viral Posts in Social Media

2024-03-02 · Yuya Sasaki, Jing Tao, Yulong Wang

Motivated by the empirical observation of power-law distributions in the credits (e.g., "likes") of viral social media posts, we introduce a high-dimensional tail index regression model and propose methods for estimation…

Limbic: Author-Based Sentiment Aspect Modeling Regularized with Word Embeddings and Discourse Relations

2018-10-01 · EMNLP 2018 10 · Zhe Zhang, Munindar Singh

We propose Limbic, an unsupervised probabilistic model that addresses the problem of discovering aspects and sentiments and associating them with authors of opinionated texts. Limbic combines three ideas, incorporating a…

DiversityGeneral ClassificationSemantic SimilaritySemantic Textual Similarity+4

Prediction-Constrained Training for Semi-Supervised Mixture and Topic Models

2017-07-23 · Michael C. Hughes, Leah Weiner, Gabriel Hope, Thomas H. McCoy Jr. 외

Supervisory signals have the potential to make low-dimensional data representations, like those learned by mixture and topic models, more interpretable and useful. We propose a framework for training latent variable mode…

PredictionSentiment AnalysisTopic Models

Graph Regularized Autoencoder and its Application in Unsupervised Anomaly Detection

2020-10-29 · Imtiaz Ahmed, Travis Galoppo, Xia Hu, Yu Ding

Dimensionality reduction is a crucial first step for many unsupervised learning tasks including anomaly detection and clustering. Autoencoder is a popular mechanism to accomplish dimensionality reduction. In order to mak…

Anomaly DetectionClusteringDimensionality ReductionUnsupervised Anomaly Detection

Survival-Supervised Topic Modeling with Anchor Words: Characterizing Pancreatitis Outcomes

2017-12-02 · George H. Chen, Jeremy C. Weiss

We introduce a new approach for topic modeling that is supervised by survival analysis. Specifically, we build on recent work on unsupervised topic modeling with so-called anchor words by providing supervision through an…

Survival Analysis