paper-with-me

홈 › Papers

Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models

2024-08-07 · Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-Yi Lee, Lama Nachman

Large Language Models (LLMs) have excelled at language understanding and generating human-level text. However, even with supervised training and human alignment, these LLMs are susceptible to adversarial attacks where malicious users can prompt the model to generate undesirable text. LLMs also inherently encode potential biases that can cause various harmful effects during interactions. Bias evaluation metrics lack standards as well as consensus and existing methods often rely on human-generated templates and annotations which are expensive and labor intensive. In this work, we train models to automatically create adversarial prompts to elicit biased responses from target LLMs. We present LLM- based bias evaluation metrics and also analyze several existing automatic evaluation methods and metrics. We analyze the various nuances of model responses, identify the strengths and weaknesses of model families, and assess where evaluation methods fall short. We compare these metrics to human evaluation and validate that the LLM-as-a-Judge metric aligns with human judgement on bias in response generation.

📄 PDF Abstract BibTeX arXiv:2408.03907

Code (0)

등록된 구현이 없습니다.

Tasks

Bias DetectionGender Bias DetectionResponse Generation

Similar Papers 제목 키워드 기반

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

2026-03-09 · Hongli Zhou, Hui Huang, Rui Zhang, Kehai Chen 외 arxiv

Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evaluating these biases is essential for ensu…

Reinforcement LearningContrastive Learning

Humans or LLMs as the Judge? A Study on Judgement Biases

2024-02-16 · Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 외

Adopting human and large language models (LLM) as judges (a.k.a human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potentia…

Misinformation

Identifying Gender Stereotypes and Biases in Automated Translation from English to Italian using Similarity Networks

2025-02-17 · Fatemeh Mohammadi, Marta Annamaria Tamborini, Paolo Ceravolo, Costanza Nardocci 외

This paper is a collaborative effort between Linguistics, Law, and Computer Science to evaluate stereotypes and biases in automated translation systems. We advocate gender-neutral translation as a means to promote gender…

Machine TranslationTranslation

Judicial Favoritism of Politicians: Evidence from Small Claims Court

2020-01-31

Multiple studies have documented racial, gender, political ideology, or ethnical biases in comparative judicial systems. Supplementing this literature, we investigate whether judges rule cases differently when one of the…

Using Item Response Theory to Measure Gender and Racial Bias of a BERT-based Automated English Speech Assessment System

2022-07-01 · NAACL (BEA) 2022 7 · Alexander Kwako, Yixin Wan, Jieyu Zhao, Kai-Wei Chang 외

Recent advances in natural language processing and transformer-based models have made it easier to implement accurate, automated English speech assessments. Yet, without careful examination, applications of these models …