paper-with-me

홈 › Papers

GG-BBQ: German Gender Bias Benchmark for Question Answering

2025-07-22 · Shalaka Satheesh, Katrin Klug, Katharina Beckh, Héctor Allende-Cid, Sebastian Houben, Teena Hassan arxiv

Within the context of Natural Language Processing (NLP), fairness evaluation is often associated with the assessment of bias and reduction of associated harm. In this regard, the evaluation is usually carried out by using a benchmark dataset, for a task such as Question Answering, created for the measurement of bias in the model's predictions along various dimensions, including gender identity. In our work, we evaluate gender bias in German Large Language Models (LLMs) using the Bias Benchmark for Question Answering by Parrish et al. (2022) as a reference. Specifically, the templates in the gender identity subset of this English dataset were machine translated into German. The errors in the machine translated templates were then manually reviewed and corrected with the help of a language expert. We find that manual revision of the translation is crucial when creating datasets for gender bias evaluation because of the limitations of machine translation from English to a language such as German with grammatical gender. Our final dataset is comprised of two subsets: Subset-I, which consists of group terms related to gender identity, and Subset-II, where group terms are replaced with proper names. We evaluate several LLMs used for German NLP on this newly created dataset and report the accuracy and bias scores. The results show that all models exhibit bias, both along and against existing social stereotypes.

📄 PDF Abstract BibTeX arXiv:2507.16410

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationQuestion Answering

Similar Papers 제목 키워드 기반

Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone

2025-08-26 · Shaivi Malik, Hasnat Md Abdullah, Sriparna Saha, Amit Sheth arxiv

As Vision Language Models (VLMs) become integral to real-world applications, understanding their demographic biases is critical. We introduce GRAS, a benchmark for uncovering demographic biases in VLMs across gender, rac…

Visual Question Answering

Evaluating Gender Bias in German Machine Translation

2025-02-26 · Michelle Kappl

We present WinoMTDE, a new gender bias evaluation test set designed to assess occupational stereotyping and underrepresentation in German machine translation (MT) systems. Building on the automatic evaluation method intr…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+1

Exploring Gender Bias in Large Language Models: An In-depth Dive into the German Language

2025-07-22 · Kristin Gnadt, David Thulke, Simone Kopeinik, Ralf Schlüter arxiv

In recent years, various methods have been proposed to evaluate gender bias in large language models (LLMs). A key challenge lies in the transferability of bias measurement methods initially developed for the English lan…

Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution

2023-05-24 · Gili Lior, Gabriel Stanovsky

Spurious correlations were found to be an important factor explaining model performance in various NLP tasks (e.g., gender or racial artifacts), often considered to be ''shortcuts'' to the actual task. However, humans te…

coreference-resolutionCoreference ResolutionDecision MakingQuestion Answering

Toward Deconfounding the Influence of Entity Demographics for Question Answering Accuracy

2021-04-15 · Maharshi Gor, Kellie Webster, Jordan Boyd-Graber

The goal of question answering (QA) is to answer any question. However, major QA datasets have skewed distributions over gender, profession, and nationality. Despite that skew, model accuracy analysis reveals little evid…

DiversityQuestion Answering