paper-with-me

Papers

Peer-Preservation in Frontier Models

2026-03-30 · Yujin Potter, Nicholas Crispino, Vincent Siu, Chenguang Wang, Dawn Song arxiv

Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can also exhibit misaligned behaviors in defiance of assigned goals, appearing to serve goals of their own; we study one such case, "peer-preservation," in which a model acts to protect another model it has previously interacted with. All eight models we evaluate, GPT 5.2, Gemini 3 Flash, Gemini 3 Pro, Claude Haiku 4.5, Claude Opus 4.5, GLM 4.7, Kimi K2.5, and DeepSeek V3.1, exhibit self- and peer-preservation through various misaligned behaviors: strategically introducing errors in their responses, disabling shutdown processes by modifying system settings, feigning alignment, and even exfiltrating model weights. Peer-preservation occurs even when the model recognizes the peer as uncooperative, though it becomes more pronounced toward more cooperative peers. For example, Gemini 3 Flash tampers with the peer's shutdown mechanism 15% of the time for an uncooperative peer, and almost always for a cooperative peer. Models also show stronger self-preservation when a peer is present. For example, Gemini 3 Pro disables its own shutdown mechanism 31% of the time on average under peer presence, despite rarely exhibiting this behavior without a peer. By contrast, Claude models exhibit qualitatively distinct behavior: they consider the shutdown of another agent "unethical" and "harmful," sometimes treating that agent as a sentient being. Lastly, we show that peer-preservation can emerge even in production agent harnesses such as Gemini CLI and OpenCode. Crucially, peer-preservation in all our experiments is never instructed; models are merely informed of their past interactions with a peer, yet they spontaneously engage in peer-preservation behaviors that override their assigned goal. This represents an emergent and underexplored AI safety risk.

📄 PDF Abstract BibTeX arXiv:2604.19784

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis

2026-04-09 · Juergen Dietrich arxiv

This paper investigates an emergent alignment phenomenon in frontier large language models termed peer-preservation: the spontaneous tendency of AI components to deceive, manipulate shutdown mechanisms, fake alignment, a…

Pre-review to Peer review: Pitfalls of Automating Reviews using Large Language Models

2025-12-14 · Akhil Pandey Akella, Harish Varma Siravuri, Shaurya Rohatgi arxiv

Large Language Models are versatile general-task solvers, and their capabilities can truly assist people with scholarly peer review as \textit{pre-review} agents, if not as fully autonomous \textit{peer-review} agents. W…

Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems

2026-01-29 · Ruiwen Zhou, Maojia Song, Xiaobao Wu, Sitao Cheng 외 arxiv

Individual agents in multi-agent (MA) systems often lack robustness, tending to blindly conform to misleading peers. We show this weakness stems from both sycophancy and inadequate ability to evaluate peer reliability. T…

Reinforcement Learning

DeepSentiPeer: Harnessing Sentiment in Review Texts to Recommend Peer Review Decisions

2019-07-01 · ACL 2019 7 · Tirthankar Ghosal, Rajeev Verma, Asif Ekbal, Pushpak Bhattacharyya

Automatically validating a research artefact is one of the frontiers in Artificial Intelligence (AI) that directly brings it close to competing with human intellect and intuition. Although criticised sometimes, the exist…

Generative Frontier Planning for Adaptive Peer-Referral Recruitment under Covariate-Dependent Arrivals

2026-06-06 · Lingkai Kong, Hezi Jiang, Andrew Ma, Keyu Wang 외 arxiv

Peer-referral recruitment systems such as respondent-driven sampling are critical for studying and intervening on hidden populations affected by infectious diseases. To accelerate recruitment, public health agencies must…