paper-with-me

Papers

Persona Setting Pitfall: Persistent Outgroup Biases in Large Language Models Arising from Social Identity Adoption

2024-09-05 · Wenchao Dong, Assem Zhunis, Dongyoung Jeong, Hyojin Chin, Jiyoung Han, Meeyoung Cha

Drawing parallels between human cognition and artificial intelligence, we explored how large language models (LLMs) internalize identities imposed by targeted prompts. Informed by Social Identity Theory, these identity assignments lead LLMs to distinguish between "we" (the ingroup) and "they" (the outgroup). This self-categorization generates both ingroup favoritism and outgroup bias. Nonetheless, existing literature has predominantly focused on ingroup favoritism, often overlooking outgroup bias, which is a fundamental source of intergroup prejudice and discrimination. Our experiment addresses this gap by demonstrating that outgroup bias manifests as strongly as ingroup favoritism. Furthermore, we successfully mitigated the inherent pro-liberal, anti-conservative bias in LLMs by guiding them to adopt the perspectives of the initially disfavored group. These results were replicated in the context of gender bias. Our findings highlight the potential to develop more equitable and balanced language models.

📄 PDF Abstract BibTeX arXiv:2409.03843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sacred or Secular? Religious Bias in AI-Generated Financial Advice

2025-03-26 · Muhammad Salar Khan, Hamza Umer

This study examines religious biases in AI-generated financial advice, focusing on ChatGPT's responses to financial queries. Using a prompt-based methodology and content analysis, we find that 50% of the financial emails…

Friction

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

2025-10-08 · Geng Liu, Feng Li, Junjie Mu, Mengxiao Zhu 외 arxiv

Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases. We investigate social identity biases in Chinese LLMs using Mandarin-s…

Generative Language Models Exhibit Social Identity Biases

2023-10-24 · Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier 외

The surge in popularity of large language models has given rise to concerns about biases that these models could learn from humans. We investigate whether ingroup solidarity and outgroup hostility, fundamental social ide…

When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents

2026-01-01 · Zongwei Wang, Bincheng Gu, Hongyu Yu, Junliang Yu 외 arxiv

This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal "us" versus "them" cues. When such group boundaries align with the agent-human …

Using Word Embeddings to Quantify Ethnic Stereotypes in 12 years of Spanish News

2021-12-01 · ALTA 2021 12 · Danielly Sorato, Diana Zavala-Rojas, Maria del Carme Colominas Ventura

The current study provides a diachronic analysis of the stereotypical portrayals concerning seven of the most prominent foreign nationalities living in Spain in a Spanish news outlet. We use 12 years (2007-2018) of news …

ArticlesWord Embeddings