paper-with-me

홈 › Papers

Regional Negative Bias in Word Embeddings Predicts Racial Animus--but only via Name Frequency

2022-01-20 · Austin Van Loon, Salvatore Giorgi, Robb Willer, Johannes Eichstaedt

The word embedding association test (WEAT) is an important method for measuring linguistic biases against social groups such as ethnic minorities in large text corpora. It does so by comparing the semantic relatedness of words prototypical of the groups (e.g., names unique to those groups) and attribute words (e.g., 'pleasant' and 'unpleasant' words). We show that anti-black WEAT estimates from geo-tagged social media data at the level of metropolitan statistical areas strongly correlate with several measures of racial animus--even when controlling for sociodemographic covariates. However, we also show that every one of these correlations is explained by a third variable: the frequency of Black names in the underlying corpora relative to White names. This occurs because word embeddings tend to group positive (negative) words and frequent (rare) words together in the estimated semantic space. As the frequency of Black names on social media is strongly correlated with Black Americans' prevalence in the population, this results in spurious anti-Black WEAT estimates wherever few Black Americans live. This suggests that research using the WEAT to measure bias should consider term frequency, and also demonstrates the potential consequences of using black-box models like word embeddings to study human cognition and behavior.

📄 PDF Abstract BibTeX arXiv:2201.08451

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeWord Embeddings

Similar Papers 제목 키워드 기반

A Transparent Framework for Evaluating Unintended Demographic Bias in Word Embeddings

2019-07-01 · ACL 2019 7 · Chris Sweeney, Maryam Najafian

Word embedding models have gained a lot of traction in the Natural Language Processing community, however, they suffer from unintended demographic biases. Most approaches to evaluate these biases rely on vector space bas…

FairnessWord Embeddings

Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation

2020-05-03 · ACL 2020 6 · Tianlu Wang, Xi Victoria Lin, Nazneen Fatema Rajani, Bryan McCann 외

Word embeddings derived from human-generated corpora inherit strong gender bias which can be further amplified by downstream models. Some commonly adopted debiasing approaches, including the seminal Hard Debias algorithm…

Word Embeddings

"Thy algorithm shalt not bear false witness": An Evaluation of Multiclass Debiasing Methods on Word Embeddings

2020-10-30 · Thalea Schlender, Gerasimos Spanakis

With the vast development and employment of artificial intelligence applications, research into the fairness of these algorithms has been increased. Specifically, in the natural language processing domain, it has been sh…

FairnessWord Embeddings

The Undesirable Dependence on Frequency of Gender Bias Metrics Based on Word Embeddings

2023-01-02 · Francisco Valentini, Germán Rosati, Diego Fernandez Slezak, Edgar Altszyler

Numerous works use word embedding-based metrics to quantify societal biases and stereotypes in texts. Recent studies have found that word embeddings can capture semantic similarity but may be affected by word frequency. …

Semantic SimilaritySemantic Textual SimilarityWord Embeddings

Improving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives

2022-06-06 · Wei Wang, Liangzhu Ge, Jingqiao Zhang, Cheng Yang

Following SimCSE, contrastive learning based methods have achieved the state-of-the-art (SOTA) performance in learning sentence embeddings. However, the unsupervised contrastive learning methods still lag far behind the …

AttributeContrastive LearningLanguage ModelingLanguage Modelling+5