paper-with-me

홈 › Papers

Bias after Prompting: Persistent Discrimination in Large Language Models

2025-09-09 · Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi, Krishna Patel, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff arxiv

A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapted models. We invalidate this assumption by studying the BTH in causal models under prompt adaptations, as prompting is an extremely popular and accessible adaptation strategy used in real-world applications. In contrast to prior work, we find that biases can transfer through prompting and that popular prompt-based mitigation methods do not consistently prevent biases from transferring. Specifically, the correlation between intrinsic biases and those after prompt adaptation remain moderate to strong across demographics and tasks -- for example, gender (rho >= 0.94) in co-reference resolution, and age (rho >= 0.98) and religion (rho >= 0.69) in question answering. Further, we find that biases remain strongly correlated when varying few-shot composition parameters, such as sample size, stereotypical content, occupational distribution and representational balance (rho >= 0.90). We evaluate several prompt-based debiasing strategies and find that different approaches have distinct strengths, but none consistently reduce bias transfer across models, tasks or demographics. These results demonstrate that correcting bias, and potentially improving reasoning ability, in intrinsic models may prevent propagation of biases to downstream tasks.

📄 PDF Abstract BibTeX arXiv:2509.08146

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Professional Presentation and Projected Power: A Case Study of Implicit Gender Information in English CVs

2022-11-17 · Jinrui Yang, Sheilla Njoto, Marc Cheong, Leah Ruppanner 외

Gender discrimination in hiring is a pertinent and persistent bias in society, and a common motivating example for exploring bias in NLP. However, the manifestation of gendered language in application materials has recei…

EMBER: Autonomous Cognitive Behaviour from Learned Spiking Neural Network Dynamics in a Hybrid LLM Architecture

2026-04-14 · William Savage arxiv

We present (Experience-Modulated Biologically-inspired Emergent Reasoning), a hybrid cognitive architecture that reorganises the relationship between large language models (LLMs) and memory: rather than augmenting an LLM…

Discrimination Against Immigrants in the Criminal Justice System: Evidence from Pretrial Detentions

2022-02-22 · Patricio Domínguez, Nicolás Grau, Damián Vergara

This paper tests for discrimination against immigrant defendants in the criminal justice system in Chile using a decade of nationwide administrative records on pretrial detentions. Observational benchmark regressions sho…

Statistical discrimination in learning agents

2021-10-21 · Edgar A. Duéñez-Guzmán, Kevin R. McKee, Yiran Mao, Ben Coppin 외

Undesired bias afflicts both human and algorithmic decision making, and may be especially prevalent when information processing trade-offs incentivize the use of heuristics. One primary example is \textit{statistical dis…

Decision MakingMulti-agent Reinforcement Learning

Can LLMs Hire Fairly? Racial Bias in Resume Screening

2026-06-27 · Zhenyu Gao, Wenxi Jiang, Yutong Yan arxiv

We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback ga…