paper-with-me

홈 › Papers

StereoSet: Measuring stereotypical bias in pretrained language models

2020-04-20 · ACL 2021 5 · Moin Nadeem, Anna Bethke, Siva Reddy

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at https://stereoset.mit.edu

📄 PDF Abstract BibTeX arXiv:2004.09456

Code (3)

moinnadeem/StereoSet 공식 구현 pytorch
kanekomasahiro/evaluate_bias_in_mlm pytorch
zalkikar/mlm-bias pytorch

Tasks

Bias DetectionMath

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Weight Decay 설명 없음
Discriminative Fine-Tuning Discriminative Fine-Tuning is a fine-tuning strategy that is used for ULMFiT type models. Instead of using the same learning rate…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
RoBERTa 설명 없음
WordPiece 설명 없음
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…

Similar Papers 제목 키워드 기반

BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla

2024-09-18 · Mahammed Kamruzzaman, Abdullah Al Monsur, Shrabon Das, Enamul Hassan 외

This study presents BanStereoSet, a dataset designed to evaluate stereotypical social biases in multilingual LLMs for the Bangla language. In an effort to extend the focus of bias research beyond English-centric datasets…

Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models

2024-08-14 · Yi-Cheng Lin, Wei-Chih Chen, Hung-Yi Lee

Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these …

Robust Evaluation Measures for Evaluating Social Biases in Masked Language Models

2024-01-21 · Yang Liu

Many evaluation measures are used to evaluate social biases in masked language models (MLMs). However, we find that these previously proposed evaluation measures are lacking robustness in scenarios with limited datasets.…

How Different Is Stereotypical Bias Across Languages?

2023-07-14 · Ibrahim Tolga Öztürk, Rostislav Nedelchev, Christian Heumann, Esteban Garces Arias 외

Recent studies have demonstrated how to assess the stereotypical bias in pre-trained English language models. In this work, we extend this branch of research in multiple different dimensions by systematically investigati…

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender

2026-05-12 · Leonor Veloso, Hinrich Schütze arxiv

Recent works have analyzed the impact of individual components of neural networks on gendered predictions, often with a focus on mitigating gender bias. However, mechanistic interpretations of gender tend to (i) focus on…