paper-with-me

Papers

Authorless Topic Models: Biasing Models Away from Known Structure

2018-08-01 · COLING 2018 8 · Laure Thompson, David Mimno

Most previous work in unsupervised semantic modeling in the presence of metadata has assumed that our goal is to make latent dimensions more correlated with metadata, but in practice the exact opposite is often true. Some users want topic models that highlight differences between, for example, authors, but others seek more subtle connections across authors. We introduce three metrics for identifying topics that are highly correlated with metadata, and demonstrate that this problem affects between 30 and 50{\%} of the topics in models trained on two real-world collections, regardless of the size of the model. We find that we can predict which words cause this phenomenon and that by selectively subsampling these words we dramatically reduce topic-metadata correlation, improve topic stability, and maintain or even improve model quality.

📄 PDF Abstract BibTeX

Code (1)

laurejt/authorless-tms 공식 구현

Tasks

Document ClassificationTopic ModelsWord Embeddings

Similar Papers 제목 키워드 기반

A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning

2022-03-22 · Hugo Berg, Siobhan Mackenzie Hall, Yash Bhalgat, Wonsuk Yang 외

Vision-language models can encode societal biases and stereotypes, but there are challenges to measuring and mitigating these multimodal harms due to lacking measurement robustness and feature degradation. To address the…

Improving Structured Text Recognition with Regular Expression Biasing

2021-11-10 · Baoguang Shi, WenFeng Cheng, Yijuan Lu, Cha Zhang 외

We study the problem of recognizing structured text, i.e. text that follows certain formats, and propose to improve the recognition accuracy of structured text by specifying regular expressions (regexes) for biasing. A b…

Decoder

Graph Convolutional Networks Meet with High Dimensionality Reduction

2019-11-07 · Mustafa Coskun

Recently, Graph Convolutional Networks (GCNs) and their variants have been receiving many research interests for learning graph-related tasks. While the GCNs have been successfully applied to this problem, some caveats i…

BenchmarkingDimensionality ReductionNode ClassificationVocal Bursts Intensity Prediction

Cross-Topic Rumor Detection using Topic-Mixtures

2021-04-01 · EACL 2021 2 · Xiaoying Ren, Jing Jiang, Ling Min Serena Khoo, Hai Leong Chieu

There has been much interest in rumor detection using deep learning models in recent years. A well-known limitation of deep learning models is that they tend to learn superficial patterns, which restricts their generaliz…

Mixture-of-Experts

Disentangling Document Topic and Author Gender in Multiple Languages: Lessons for Adversarial Debiasing

2021-04-01 · EACL (WASSA) 2021 4 · Erenay Dayanik, Sebastian Padó

Text classification is a central tool in NLP. However, when the target classes are strongly correlated with other textual attributes, text classification models can pick up “wrong” features, leading to bad generalization…

Classificationtext-classificationText Classification