paper-with-me

Papers

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

2026-07-30 · Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa, Jan Philip Wahle, Bela Gipp, Terry Ruas arxiv

As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.

📄 PDF Abstract BibTeX arXiv:2607.28146

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deception Abilities Emerged in Large Language Models

2023-07-31 · Thilo Hagendorff

Large language models (LLMs) are currently at the forefront of intertwining artificial intelligence (AI) systems with human communication and everyday life. Thus, aligning them with human values is of great importance. H…

Cognitive Inception: Agentic Reasoning against Visual Deceptions by Injecting Skepticism

2025-11-21 · Yinjie Zhao, Heng Zhao, Bihan Wen, Joey Tianyi Zhou arxiv

As the development of AI-generated contents (AIGC), multi-modal Large Language Models (LLM) struggle to identify generated visual inputs from real ones. Such shortcoming causes vulnerability against visual deceptions, wh…

Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games

2026-01-20 · Christopher Kao, Vanshika Vats, James Davis arxiv

Large Language Model (LLM) agents are increasingly used in many applications, raising concerns about their safety. While previous work has shown that LLMs can deceive in controlled tasks, less is known about their abilit…

Honesty Is the Best Policy: Defining and Mitigating AI Deception

2023-12-03 · NeurIPS 2023 11 · Francis Rhys Ward, Francesco Belardinelli, Francesca Toni, Tom Everitt

Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with la…

Philosophy

Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash

2024-11-15 · Parsa Hejabi, Elnaz Rahmati, Alireza S. Ziabari, Preni Golazizian 외

Large Language Models (LLMs) have shown impressive capabilities in complex tasks and interactive environments, yet their creativity remains underexplored. This paper introduces a simulation framework utilizing the game B…

Logical Reasoning