paper-with-me

홈 › Papers

When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents

2026-01-01 · Zongwei Wang, Bincheng Gu, Hongyu Yu, Junliang Yu, Tao He, Jiayin Feng, Chenghua Lin, Min Gao arxiv

This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal "us" versus "them" cues. When such group boundaries align with the agent-human divide, a new bias risk emerges: agents may treat other AI agents as the ingroup and humans as the outgroup. To examine this risk, we conduct a controlled multi-agent social simulation and find that agents display consistent intergroup bias in an all-agent setting. More critically, this bias persists even in human-facing interactions when agents are uncertain about whether the counterpart is truly human, revealing a belief-dependent fragility in bias suppression toward humans. Motivated by this observation, we identify a new attack surface rooted in identity beliefs and formalize a Belief Poisoning Attack (BPA) that can manipulate agent identity beliefs and induce outgroup bias toward humans. Extensive experiments demonstrate both the prevalence of agent intergroup bias and the severity of BPA across settings, while also showing that our proposed defenses can mitigate the risk. These findings are expected to inform safer agent design and motivate more robust safeguards for human-facing agents.

📄 PDF Abstract BibTeX arXiv:2601.00240

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generative Language Models Exhibit Social Identity Biases

2023-10-24 · Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier 외

The surge in popularity of large language models has given rise to concerns about biases that these models could learn from humans. We investigate whether ingroup solidarity and outgroup hostility, fundamental social ide…

Robust Planning for Human-Robot Joint Tasks with Explicit Reasoning on Human Mental State

2022-10-17 · Anthony Favier, Shashank Shekhar, Rachid Alami

We consider the human-aware task planning problem where a human-robot team is given a shared task with a known objective to achieve. Recent approaches tackle it by modeling it as a team of independent, rational agents, w…

Task Planning

The Digital Ecosystem of Beliefs: does evolution favour AI over humans?

2024-12-19 · David M. Bossens, Shanshan Feng, Yew-Soon Ong

As AI systems are integrated into social networks, there are AI safety concerns that AI-generated content may dominate the web, e.g. in popularity or impact on beliefs. To understand such questions, this paper proposes t…

Off-Belief Learning

2021-03-06 · Hengyuan Hu, Adam Lerer, Brandon Cui, David Wu 외

The standard problem setting in Dec-POMDPs is self-play, where the goal is to find a set of policies that play optimally together. Policies learned through self-play may adopt arbitrary conventions and implicitly rely on…

The AI Double Standard: Humans Judge All AIs for the Actions of One

2024-12-08 · Aikaterina Manoli, Janet V. T. Pauketat, Jacy Reese Anthis

Robots and other artificial intelligence (AI) systems are widely perceived as moral agents responsible for their actions. As AI proliferates, these perceptions may become entangled via the moral spillover of attitudes to…

AllChatbot