Choose Your Lenses: Flaws in Gender Bias Evaluation
Considerable efforts to measure and mitigate gender bias in recent years have led to the introduction of an abundance of tasks, datasets, and metrics used in this vein. In this position paper, we assess the current paradigm of gender bias evaluation and identify several flaws in it. First, we highlight the importance of extrinsic bias metrics that measure how a model's performance on some task is affected by gender, as opposed to intrinsic evaluations of model representations, which are less strongly connected to specific harms to people interacting with systems. We find that only a few extrinsic metrics are measured in most studies, although more can be measured. Second, we find that datasets and metrics are often coupled, and discuss how their coupling hinders the ability to obtain reliable conclusions, and how one may decouple them. We then investigate how the choice of the dataset and its composition, as well as the choice of the metric, affect bias measurement, finding significant variations across each of them. Finally, we propose several guidelines for more reliable gender bias evaluation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Indigenous Language Revitalization and the Dilemma of Gender Bias
Natural Language Processing (NLP), through its several applications, has been considered as one of the most valuable field in interdisciplinary researches, as well as in computer science. However, it is not without its f…
Word EmbeddingsColombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
One of the goals of fairness research in NLP is to measure and mitigate stereotypical biases that are propagated by NLP systems. However, such work tends to focus on single axes of bias (most often gender) and the Englis…
FairnessGender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
Large language models (LLMs) acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. Prior studies have demonstrated model generations favor one gender or exhi…
Decision MakingWhat is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models
Bias is a disproportionate prejudice in favor of one side against another. Due to the success of transformer-based Masked Language Models (MLMs) and their impact on many NLP tasks, a systematic evaluation of bias in thes…
SentenceInvestigating the Roots of Gender Bias in Machine Translation: Observations on Gender Transfer between French and English
This paper aims at identifying the inner mechanisms that make a translation model choose a masculine rather than a feminine form, an essential step to mitigate gender bias in MT. We conduct two series of experiments u…
Machine TranslationTranslation