paper-with-me

Papers

On Measures of Biases and Harms in NLP

2021-08-07 · Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, Kai-Wei Chang

Recent studies show that Natural Language Processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality. To create interventions and mitigate these biases and associated harms, it is vital to be able to detect and measure such biases. While existing works propose bias evaluation and mitigation methods for various tasks, there remains a need to cohesively understand the biases and the specific harms they measure, and how different measures compare with each other. To address this gap, this work presents a practical framework of harms and a series of questions that practitioners can answer to guide the development of bias measures. As a validation of our framework and documentation questions, we also present several case studies of how existing bias measures in NLP -- both intrinsic measures of bias in representations and extrinsic measures of bias of downstream applications -- can be aligned with different harms and how our proposed documentation questions facilitates more holistic understanding of what bias measures are measuring.

📄 PDF Abstract BibTeX arXiv:2108.03362

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring Political Bias in Large Language Models: What Is Said and How It Is Said

2024-03-27 · Yejin Bang, Delong Chen, Nayeon Lee, Pascale Fung

We propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues. Existing benchmarks and measures focus on gender and racial biases. However, pol…

A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning

2022-03-22 · Hugo Berg, Siobhan Mackenzie Hall, Yash Bhalgat, Wonsuk Yang 외

Vision-language models can encode societal biases and stereotypes, but there are challenges to measuring and mitigating these multimodal harms due to lacking measurement robustness and feature degradation. To address the…

From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards

2024-03-20 · Khaoula Chehbouni, Megha Roshan, Emmanuel Ma, Futian Andrew Wei 외

Recent progress in large language models (LLMs) has led to their widespread adoption in various domains. However, these advancements have also introduced additional safety risks and raised concerns regarding their detrim…

Safe Reinforcement Learning

Measuring What Matters: Connecting AI Ethics Evaluations to System Attributes, Hazards, and Harms

2025-10-11 · Shalaleh Rismani, Renee Shelby, Leah Davis, Negar Rostamzadeh 외 arxiv

Over the past decade, an ecosystem of measures has emerged to evaluate the social and ethical implications of AI systems, largely shaped by high-level ethics principles. These measures are developed and used in fragmente…

"Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters

2023-10-13 · Yixin Wan, George Pu, Jiao Sun, Aparna Garimella 외

Large Language Models (LLMs) have recently emerged as an effective tool to assist individuals in writing various types of content, including professional documents such as recommendation letters. Though bringing convenie…

BenchmarkingFairnessHallucination