paper-with-me

Papers

Collecting a Large-Scale Gender Bias Dataset for Coreference Resolution and Machine Translation

2021-09-08 · Findings (EMNLP) 2021 11 · Shahar Levy, Koren Lazar, Gabriel Stanovsky

Recent works have found evidence of gender bias in models of machine translation and coreference resolution using mostly synthetic diagnostic datasets. While these quantify bias in a controlled experiment, they often do so on a small scale and consist mostly of artificial, out-of-distribution sentences. In this work, we find grammatical patterns indicating stereotypical and non-stereotypical gender-role assignments (e.g., female nurses versus male dancers) in corpora from three domains, resulting in a first large-scale gender bias dataset of 108K diverse real-world English sentences. We manually verify the quality of our corpus and use it to evaluate gender bias in various coreference resolution and machine translation models. We find that all tested models tend to over-rely on gender stereotypes when presented with natural inputs, which may be especially harmful when deployed in commercial systems. Finally, we show that our dataset lends itself to finetuning a coreference resolution model, finding it mitigates bias on a held out set. Our dataset and models are publicly available at www.github.com/SLAB-NLP/BUG. We hope they will spur future research into gender bias evaluation mitigation techniques in realistic settings.

📄 PDF Abstract BibTeX arXiv:2109.03858

Code (1)

slab-nlp/bug 공식 구현

Tasks

coreference-resolutionCoreference ResolutionDiagnosticMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Gender Bias in Text: Labeled Datasets and Lexicons

2022-01-21 · Jad Doughman, Wael Khreich

Language has a profound impact on our thoughts, perceptions, and conceptions of gender roles. Gender-inclusive language is, therefore, a key tool to promote social inclusion and contribute to achieving gender equality. C…

Unveiling Gender Bias in Terms of Profession Across LLMs: Analyzing and Addressing Sociological Implications

2023-07-18 · Vishesh Thakur

Gender bias in artificial intelligence (AI) and natural language processing has garnered significant attention due to its potential impact on societal perceptions and biases. This research paper aims to analyze gender bi…

Data Augmentation

Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens

2025-11-05 · Hellina Hailu Nigatu, Bethelhem Yemane Mamo, Bontu Fufa Balcha, Debora Taye Tesfaye 외 arxiv

As low-resourced languages are increasingly incorporated into NLP research, there is an emphasis on collecting large-scale datasets. But in prioritizing quantity over quality, we risk 1) building language technologies th…

Machine Translation

Probing Intersectional Biases in Vision-Language Models with Counterfactual Examples

2023-10-04 · Phillip Howard, Avinash Madasu, Tiep Le, Gustavo Lujan Moreno 외

While vision-language models (VLMs) have achieved remarkable performance improvements recently, there is growing evidence that these models also posses harmful biases with respect to social attributes such as gender and …

counterfactual

SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples

2023-11-30 · CVPR 2024 1 · Phillip Howard, Avinash Madasu, Tiep Le, Gustavo Lujan Moreno 외

While vision-language models (VLMs) have achieved remarkable performance improvements recently, there is growing evidence that these models also posses harmful biases with respect to social attributes such as gender and …

counterfactual