Authorless Topic Models: Biasing Models Away from Known Structure
Most previous work in unsupervised semantic modeling in the presence of metadata has assumed that our goal is to make latent dimensions more correlated with metadata, but in practice the exact opposite is often true. Some users want topic models that highlight differences between, for example, authors, but others seek more subtle connections across authors. We introduce three metrics for identifying topics that are highly correlated with metadata, and demonstrate that this problem affects between 30 and 50{\%} of the topics in models trained on two real-world collections, regardless of the size of the model. We find that we can predict which words cause this phenomenon and that by selectively subsampling these words we dramatically reduce topic-metadata correlation, improve topic stability, and maintain or even improve model quality.
Code (1)
Tasks
Document ClassificationTopic ModelsWord EmbeddingsSimilar Papers 제목 키워드 기반
A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning
Vision-language models can encode societal biases and stereotypes, but there are challenges to measuring and mitigating these multimodal harms due to lacking measurement robustness and feature degradation. To address the…
Improving Structured Text Recognition with Regular Expression Biasing
We study the problem of recognizing structured text, i.e. text that follows certain formats, and propose to improve the recognition accuracy of structured text by specifying regular expressions (regexes) for biasing. A b…
DecoderGraph Convolutional Networks Meet with High Dimensionality Reduction
Recently, Graph Convolutional Networks (GCNs) and their variants have been receiving many research interests for learning graph-related tasks. While the GCNs have been successfully applied to this problem, some caveats i…
BenchmarkingDimensionality ReductionNode ClassificationVocal Bursts Intensity PredictionCross-Topic Rumor Detection using Topic-Mixtures
There has been much interest in rumor detection using deep learning models in recent years. A well-known limitation of deep learning models is that they tend to learn superficial patterns, which restricts their generaliz…
Mixture-of-ExpertsDisentangling Document Topic and Author Gender in Multiple Languages: Lessons for Adversarial Debiasing
Text classification is a central tool in NLP. However, when the target classes are strongly correlated with other textual attributes, text classification models can pick up “wrong” features, leading to bad generalization…
Classificationtext-classificationText Classification