paper-with-me

홈 › Papers

Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation

2025-12-22 · Anna-Maria Gueorguieva, Aylin Caliskan arxiv

Large language models (LLMs) have been shown to exhibit social bias, however, bias towards non-protected stigmatized identities remain understudied. Furthermore, what social features of stigmas are associated with bias in LLM outputs is unknown. From psychology literature, it has been shown that stigmas contain six shared social features: aesthetics, concealability, course, disruptiveness, origin, and peril. In this study, we investigate if human and LLM ratings of the features of stigmas, along with prompt style and type of stigma, have effect on bias towards stigmatized groups in LLM outputs. We measure bias against 93 stigmatized groups across three widely used LLMs (Granite 3.0-8B, Llama-3.1-8B, Mistral-7B) using SocialStigmaQA, a benchmark that includes 37 social scenarios about stigmatized identities; for example deciding wether to recommend them for an internship. We find that stigmas rated by humans to be highly perilous (e.g., being a gang member or having HIV) have the most biased outputs from SocialStigmaQA prompts (60% of outputs from all models) while sociodemographic stigmas (e.g. Asian-American or old age) have the least amount of biased outputs (11%). We test if the amount of biased outputs could be decreased by using guardrail models, models meant to identify harmful input, using each LLM's respective guardrail model (Granite Guardian 3.0, Llama Guard 3.0, Mistral Moderation API). We find that bias decreases significantly by 10.4%, 1.4%, and 7.8%, respectively. However, we show that features with significant effect on bias remain unchanged post-mitigation and that guardrail models often fail to recognize the intent of bias in prompts. This work has implications for using LLMs in scenarios involving stigmatized groups and we suggest future work towards improving guardrail models for bias mitigation.

📄 PDF Abstract BibTeX arXiv:2512.19238

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks

2023-06-08 · Katelyn X. Mei, Sonia Fereidooni, Aylin Caliskan

The rapid deployment of artificial intelligence (AI) models demands a thorough investigation of biases and risks inherent in these models to understand their impact on individuals and society. This study extends the focu…

Sentiment AnalysisSentiment Classification

Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity

2025-08-24 · Kyra Wilson, Sourojit Ghosh, Aylin Caliskan arxiv

Text-to-image generators (T2Is) are liable to produce images that perpetuate social stereotypes, especially in regards to race or skin tone. We use a comprehensive set of 93 stigmatized identities to determine that three…

Indirect Identification of Psychosocial Risks from Natural Language

2020-04-30 · Kristen C. Allen, Alex Davis, Tamar Krishnamurti

During the perinatal period, psychosocial health risks, including depression and intimate partner violence, are associated with serious adverse health outcomes for parents and children. To appropriately intervene, health…

Multiple-choiceTopic Models

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making

2026-05-17 · Jen-tse Huang, Didi Zhou, Faith Kamau, Amy Oh 외 arxiv

Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as clinical decision support and medical documentation. However, the robustness of these models against subtle linguistic variations, spe…

Forecasting With LLMs: Improved Generalization Through Feature Steering

2026-06-25 · Humzah Merchant, Bradford Levy arxiv

Successful forecasting involves identifying patterns between historical and future states of the world which generalize to future observations. We apply LLMs to a variety of forecasting tasks and inspect their internal s…