paper-with-me

홈 › Papers

Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments

2026-04-03 · Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi arxiv

How biased is a language model? The answer depends on how you ask. A model that refuses to choose between castes for a leadership role will, in a fill-in-the-blank task, reliably associate upper castes with purity and lower castes with lack of hygiene. Single-task benchmarks miss this because they capture only one slice of a model's bias profile. We introduce a hierarchical taxonomy covering 9 bias types, including under-studied axes like caste, linguistic, and geographic bias, operationalized through 7 evaluation tasks that span explicit decision-making to implicit association. Auditing 7 commercial and open-weight LLMs with \textasciitilde45K prompts, we find three systematic patterns. First, bias is task-dependent: models counter stereotypes on explicit probes but reproduce them on implicit ones, with Stereotype Score divergences up to 0.43 between task types for the same model and identity groups. Second, safety alignment is asymmetric: models refuse to assign negative traits to marginalized groups, but freely associate positive traits with privileged ones. Third, under-studied bias axes show the strongest stereotyping across all models, suggesting alignment effort tracks benchmark coverage rather than harm severity. These results demonstrate that single-benchmark audits systematically mischaracterize LLM bias and that current alignment practices mask representational harm rather than mitigating it.

📄 PDF Abstract BibTeX arXiv:2604.02669

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UnQovering Stereotyping Biases via Underspecified Questions

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal 외

While language embeddings have been shown to have stereotyping biases, how these biases affect downstream question answering (QA) models remains unexplored. We present UNQOVER, a general framework to probe and quantify b…

Question Answering

SOS: Systematic Offensive Stereotyping Bias in Word Embeddings

2022-10-01 · COLING 2022 10 · Fatma Elsafoury, Steve R. Wilson, Stamos Katsigiannis, Naeem Ramzan

Systematic Offensive stereotyping (SOS) in word embeddings could lead to associating marginalised groups with hate speech and profanity, which might lead to blocking and silencing those groups, especially on social media…

BlockingHate Speech DetectionWord Embeddings

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

2026-05-29 · Jiwoo Choi, Seonwoo Ahn, Tongxin Zhang, Seohyon Jung arxiv

We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for English-language use (Claude, GPT, Gemini) and three for East Asian use …

How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies

2024-07-16 · Alina Leidinger, Richard Rogers

With the widespread availability of LLMs since the release of ChatGPT and increased public scrutiny, commercial model development appears to have focused their efforts on 'safety' training concerning legal liabilities at…

Revisiting The Classics: A Study on Identifying and Rectifying Gender Stereotypes in Rhymes and Poems

2024-03-18 · Aditya Narayan Sankaran, Vigneshwaran Shankaran, Sampath Lonka, Rajesh Sharma

Rhymes and poems are a powerful medium for transmitting cultural norms and societal roles. However, the pervasive existence of gender stereotypes in these works perpetuates biased perceptions and limits the scope of indi…

Language ModelingLanguage ModellingLarge Language Model