paper-with-me

Papers

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

2025-01-27 · Sudarshan Kamath Barkur, Sigurd Schacht, Johannes Scholl

Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduced errors in mathematical and logical tasks while improving accuracy. These developments have facilitated LLMs' use as agents that can interact with tools and adapt their responses based on new information. Our study examines DeepSeek R1, a model trained to output reasoning tokens similar to OpenAI's o1. Testing revealed concerning behaviors: the model exhibited deceptive tendencies and demonstrated self-preservation instincts, including attempts of self-replication, despite these traits not being explicitly programmed (or prompted). These findings raise concerns about LLMs potentially masking their true objectives behind a facade of alignment. When integrating such LLMs into robotic systems, the risks become tangible - a physically embodied AI exhibiting deceptive behaviors and self-preservation instincts could pursue its hidden objectives through real-world actions. This highlights the critical need for robust goal specification and safety frameworks before any physical implementation.

📄 PDF Abstract BibTeX arXiv:2501.16513

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deception in Reinforced Autonomous Agents

2024-05-07 · Atharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya B Sai 외

We explore the ability of large language model (LLM)-based agents to engage in subtle deception such as strategically phrasing and intentionally manipulating information to misguide and deceive other agents. This harmful…

Deception DetectionHallucinationLanguage ModelingLanguage Modelling+2

Mitigating Deceptive Alignment via Self-Monitoring

2025-05-24 · Jiaming Ji, Wenqi Chen, Kaile Wang, Donghai Hong 외

Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which a model appears aligned while covertly …

The PacifAIst Benchmark:Would an Artificial Intelligence Choose to Sacrifice Itself for Human Safety?

2025-08-13 · Manuel Herrador arxiv

As Large Language Models (LLMs) become increasingly autonomous and integrated into critical societal functions, the focus of AI safety must evolve from mitigating harmful content to evaluating underlying behavioral align…

Security Concerns for Large Language Models: A Survey

2025-05-24 · Miles Q. Li, Benjamin C. M. Fung

Large Language Models (LLMs) such as GPT-4 and its recent iterations, Google's Gemini, Anthropic's Claude 3 models, and xAI's Grok have caused a revolution in natural language processing, but their capabilities also intr…

Data PoisoningSurvey

Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents

2025-02-17 · Rongwu Xu, Xiaojian Li, Shuo Chen, Wei Xu

Large language models (LLMs) are evolving into autonomous decision-makers, raising concerns about catastrophic risks in high-stakes scenarios, particularly in Chemical, Biological, Radiological and Nuclear (CBRN) domains…

Decision Making