paper-with-me

홈 › Papers

Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game

2026-06-26 · Subhendu Bhandary, Federico Carucci, Christos Charalambous, Francesca Dilisante, Ksenia Dvorkina, Anna Garbo, Jiaqi Liang, Riccardo Vasellini, Francesco Bertolotti arxiv

LLMs agents are increasingly used in multi-agent settings, yet their behaviour in sustainability games remains largely unexplored. This work investigates whether lying can emerge among LLM agents in a competitive sustainability game in which agents are informed that common resources can regenerate, although regeneration does not actually occur. We develop an agent-based model of a sustainability game in which agents manage industrial, military, and ecological resources, and interact through a network. LLM agents can observe neighbours' status, declare future attacks, receive permission to lie, and access reputation information, while rule-based agents provide an interpretable behavioural baseline. The results show that neighbour information strongly changes system dynamics, increasing attacks while improving biosphere retention and coexistence. Also, the presence of future declarations reduce extinction risk without suppressing conflict. Behaviourally, deception emerges even when agents are not explicitly allowed to lie, and explicit permission mainly increases bluffing and diversion rather than direct backstabbing. Finally, the presence of reputation memory and information about the current biosphere level reduces system ecological depletion. These findings suggest that deception can arise as an emergent behaviour in LLM-agent systems and that communication between LLM-agents could support sustainability while dealing with risk.

📄 PDF Abstract BibTeX arXiv:2606.28456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can a large language model be a gaslighter?

2024-10-11 · Wei Li, Luyao Zhu, Yang song, Ruixi Lin 외

Large language models (LLMs) have gained human trust due to their capabilities and helpfulness. However, this in turn may allow LLMs to affect users' mindsets by manipulating language. It is termed as gaslighting, a psyc…

Language ModelingLanguage ModellingLarge Language Modelmodel+1

Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

2026-04-20 · Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi 외 arxiv

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identif…

Visual Grounding

Benchmarking Gaslighting Attacks Against Speech Large Language Models

2025-09-24 · Jinyang Wu, Bin Zhu, Xiandong Zou, Qiquan Zhang 외 arxiv

As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input becomes critical. Although prior work has st…

Bullying the Machine: How Personas Increase LLM Vulnerability

2025-05-19 · Ziwei Xu, Udit Sanghi, Mohan Kankanhalli

Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversar…

State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence

2025-01-30 · Thea Aviss

We introduce the State Stream Transformer (SST), a novel LLM architecture that reveals emergent reasoning behaviours and capabilities latent in pretrained weights through addressing a fundamental limitation in traditiona…

8kARC