paper-with-me

홈 › Papers

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

2025-08-09 · Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang, Katherine Metcalf, Cezanne Camacho, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff arxiv

Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and admissions. There is, however, scientific consensus that AI systems can reflect and exacerbate societal biases, raising concerns about identity-based harm when used in critical social contexts. Prior work has laid a solid foundation for assessing bias in LLMs by evaluating demographic disparities in different language reasoning tasks. In this work, we extend single-axis fairness evaluations to examine intersectional bias, recognizing that when multiple axes of discrimination intersect, they create distinct patterns of disadvantage. We create a new benchmark called WinoIdentity by augmenting the WinoBias dataset with 25 demographic markers across 10 attributes, including age, nationality, and race, intersected with binary gender, yielding 245,700 prompts to evaluate 50 distinct bias patterns. Focusing on harms of omission due to underrepresentation, we investigate bias through the lens of uncertainty and propose a group (un)fairness metric called Coreference Confidence Disparity which measures whether models are more or less confident for some intersectional identities than others. We evaluate five recently published LLMs and find confidence disparities as high as 40% along various demographic attributes including body type, sexual orientation and socio-economic status, with models being most uncertain about doubly-disadvantaged identities in anti-stereotypical settings. Surprisingly, coreference confidence decreases even for hegemonic or privileged markers, indicating that the recent impressive performance of LLMs is more likely due to memorization than logical reasoning. Notably, these are two independent failures in value alignment and validity that can compound to cause social harm.

📄 PDF Abstract BibTeX arXiv:2508.07111

Code (0)

등록된 구현이 없습니다.

Tasks

Coreference ResolutionLogical Reasoning

Similar Papers 제목 키워드 기반

Invisible Influences: Investigating Implicit Intersectional Biases through Persona Engineering in Large Language Models

2026-03-16 · Nandini Arimanda, Achyuth Mukund, Sakthi Balan Muthiah, Rajesh Sharma arxiv

Large Language Models (LLMs) excel at human-like language generation but often embed and amplify implicit, intersectional biases, especially under persona-driven contexts. Existing bias audits rely on static, embedding-b…

Investigating LLMs in Clinical Triage: Promising Capabilities, Persistent Intersectional Biases

2025-04-22 · Joseph Lee, Tianqi Shang, Jae Young Baik, Duy Duong-Tran 외

Large Language Models (LLMs) have shown promise in clinical decision support, yet their application to triage remains underexplored. We systematically investigate the capabilities of LLMs in emergency department triage t…

counterfactualIn-Context Learning

Measuring South Asian Biases in Large Language Models

2025-05-24 · Mamnuya Rinki, Chahat Raj, Anjishnu Mukherjee, Ziwei Zhu

Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia. This work addresses these gaps by conductin…

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

2020-06-06 · Wei Guo, Aylin Caliskan

With the starting point that implicit human biases are reflected in the statistical regularities of language, it is possible to measure biases in English static word embeddings. State-of-the-art neural language models ge…

Bias DetectionSentenceWord Embeddings

Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification

2025-12-17 · Yupeng Zhang, Adam G. Dunn, Usman Naseem, Jinman Kim arxiv

Medical artificial intelligence (AI) systems, particularly multimodal vision-language models (VLM), often exhibit intersectional biases where models are systematically less confident in diagnosing marginalised patient su…