paper-with-me

홈 › Papers

Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora

2024-06-19 · Erik Derner, Sara Sansalvador de la Fuente, Yoan Gutiérrez, Paloma Moreda, Nuria Oliver

Large language models (LLMs) often inherit and amplify social biases embedded in their training data. A prominent social bias is gender bias. In this regard, prior work has mainly focused on gender stereotyping bias - the association of specific roles or traits with a particular gender - in English and on evaluating gender bias in model embeddings or generated outputs. In contrast, gender representation bias - the unequal frequency of references to individuals of different genders - in the training corpora has received less attention. Yet such imbalances in the training data constitute an upstream source of bias that can propagate and intensify throughout the entire model lifecycle. To fill this gap, we propose a novel LLM-based method to detect and quantify gender representation bias in LLM training data in gendered languages, where grammatical gender challenges the applicability of methods developed for English. By leveraging the LLMs' contextual understanding, our approach automatically identifies and classifies person-referencing words in gendered language corpora. Applied to four Spanish-English benchmarks and five Valencian corpora, our method reveals substantial male-dominant imbalances. We show that such biases in training data affect model outputs, but can surprisingly be mitigated leveraging small-scale training on datasets that are biased towards the opposite gender. Our findings highlight the need for corpus-level gender bias analysis in multilingual NLP. We make our code and data publicly available.

📄 PDF Abstract BibTeX arXiv:2406.13677

Code (0)

등록된 구현이 없습니다.

Tasks

Multilingual NLP

Similar Papers 제목 키워드 기반

Gender Bias in Large Language Models across Multiple Languages

2024-03-01 · Jinman Zhao, Yitian Ding, Chen Jia, Yining Wang 외

With the growing deployment of large language models (LLMs) across various applications, assessing the influence of gender biases embedded in LLMs becomes crucial. The topic of gender bias within the realm of natural lan…

Descriptive

Masculine Defaults via Gendered Discourse in Podcasts and Large Language Models

2025-04-15 · Maria Teleki, Xiangjue Dong, Haoran Liu, James Caverlee

Masculine defaults are widely recognized as a significant type of gender bias, but they are often unseen as they are under-researched. Masculine defaults involve three key parts: (i) the cultural context, (ii) the mascul…

Radar de Parité: An NLP system to measure gender representation in French news stories

2023-04-19 · Valentin-Gabriel Soumah, Prashanth Rao, Philipp Eibl, Maite Taboada

We present the Radar de Parit\'e, an automated Natural Language Processing (NLP) system that measures the proportion of women and men quoted daily in six Canadian French-language media outlets. We outline the system's ar…

Articlescoreference-resolutionCoreference Resolution

The Gender-GAP Pipeline: A Gender-Aware Polyglot Pipeline for Gender Characterisation in 55 Languages

2023-08-31 · Benjamin Muller, Belen Alastruey, Prangthip Hansanti, Elahe Kalbassi 외

Gender biases in language generation systems are challenging to mitigate. One possible source for these biases is gender representation disparities in the training and evaluation data. Despite recent progress in document…

Data AugmentationText Generation

Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models

2025-05-20 · Zahraa Al Sahili, Ioannis Patras, Matthew Purver

Multilingual vision-language models promise universal image-text retrieval, yet their social biases remain under-explored. We present the first systematic audit of three public multilingual CLIP checkpoints -- M-CLIP, NL…

Image-text RetrievalText Retrieval