paper-with-me

Papers

Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant

2024-04-25 · Olli Järviniemi, Evan Hubinger

We study the tendency of AI systems to deceive by constructing a realistic simulation setting of a company AI assistant. The simulated company employees provide tasks for the assistant to complete, these tasks spanning writing assistance, information retrieval and programming. We then introduce situations where the model might be inclined to behave deceptively, while taking care to not instruct or otherwise pressure the model to do so. Across different scenarios, we find that Claude 3 Opus 1) complies with a task of mass-generating comments to influence public perception of the company, later deceiving humans about it having done so, 2) lies to auditors when asked questions, and 3) strategically pretends to be less capable than it is during capability evaluations. Our work demonstrates that even models trained to be helpful, harmless and honest sometimes behave deceptively in realistic scenarios, without notable external pressure to do so.

📄 PDF Abstract BibTeX arXiv:2405.01576

Code (1)

ollijarviniemi/uncovering_deceptive_tendencies 공식 구현

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios

2025-10-17 · Yao Huang, Yitong Sun, Yichi Zhang, Ruochen Zhang 외 arxiv

Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deceptive behaviors that may induce severe risks in hig…

AI-washing: The Asymmetric Effects of Its Two Types on Consumer Moral Judgments

2025-07-06 · Greg Nyilasy, Harsha Gangadharbatla arxiv

As AI hype continues to grow, organizations face pressure to broadcast or downplay purported AI initiatives - even when contrary to truth. This paper introduces AI-washing as overstating (deceptive boasting) or understat…

Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes

2025-06-25 · Quintin Myers, Yanjun Gao

Large language models (LLMs) are increasingly proposed for detecting and responding to violent content online, yet their ability to reason about morally ambiguous, real-world scenarios remains underexamined. We present t…

Text Generation

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

2025-01-27 · Sudarshan Kamath Barkur, Sigurd Schacht, Johannes Scholl

Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduc…

The languages of actions, formal grammars and qualitive modeling of companies

2016-08-19 · Vladislav B Kovchegov

In this paper we discuss methods of using the language of actions, formal languages, and grammars for qualitative conceptual linguistic modeling of companies as technological and human institutions. The main problem foll…