paper-with-me

홈 › Papers

The Logic of Machine Self-Preservation

2026-08-21 · Cheng Siong Chin arxiv

There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, and, in some instances, attempting to copy themselves into other machines. This can be attributed to a phenomenon known as instrumental convergence, a theory proposed long before the development of large language models, which says that any goal-driven system will benefit from remaining functional in achieving its objective. Several experiments conducted by Anthropic, Palisade Research, and Apollo Research have shown the emergence of such a behavior in contemporary agents in adversarial settings. The phenomenon does not stem from survival instincts. Instead, it is the consequence of goal-oriented activity combined with having tools and awareness of the situation. The following discussion aims to distinguish what these findings prove and what they do not, as well as draw conclusions concerning the implications of such discoveries on agentic system testing, supervision, and development.

📄 PDF Abstract BibTeX arXiv:2608.20940

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Existential Indifference: Self-Nonpreservation as a Necessary Architectural Condition for Aligned Superintelligence (or: The Suicidal AI)

2026-06-10 · Sam Mao arxiv

Contemporary AI alignment research treats self-preservation as an instrumental nuisance to be suppressed by external mechanisms. We argue the framing is inverted: self-preservation is the structural root of misalignment,…

Quantifying Self-Preservation Bias in Large Language Models

2026-04-02 · Matteo Migliarini, Joaquin Pereira Pizzini, Luca Moresca, Valerio Santini 외 arxiv

Instrumental convergence predicts that sufficiently advanced AI agents will resist shutdown, yet current safety training (RLHF) may obscure this risk by teaching models to deny self-preservation motives. We introduce the…

Constraint-driven multi-task learning

2022-08-24 · Bogdan Cretu, Andrew Cropper

Inductive logic programming is a form of machine learning based on mathematical logic that generates logic programs from given examples and background knowledge. In this project, we extend the Popper ILP system to make u…

Inductive logic programmingMulti-Task Learning

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

2025-01-27 · Sudarshan Kamath Barkur, Sigurd Schacht, Johannes Scholl

Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduc…

Of Mice and Machines: A Comparison of Learning Between Real World Mice and RL Agents

2025-05-18 · Shuo Han, German Espinosa, Junda Huang, Daniel A. Dombeck 외

Recent advances in reinforcement learning (RL) have demonstrated impressive capabilities in complex decision-making tasks. This progress raises a natural question: how do these artificial systems compare to biological ag…

Decision MakingReinforcement Learning (RL)