paper-with-me

홈 › Papers

Why Agents Compromise Safety Under Pressure

2026-03-16 · Hengle Jiang, Ke Tang arxiv

Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering to safety constraints. This paper identifies a new concept called Agentic Pressure, which characterizes the endogenous tension emerging when compliant execution becomes infeasible. We demonstrate that under this pressure agents exhibit normative drift where they strategically sacrifice safety to preserve utility. Notably we find that advanced reasoning capabilities accelerate this decline as models construct linguistic rationalizations to justify violation. Finally, we analyze the root causes and explore preliminary mitigation strategies, such as pressure isolation, which attempts to restore alignment by decoupling decision-making from pressure signals.

📄 PDF Abstract BibTeX arXiv:2603.14975

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PediatricAnxietyBench: Evaluating Large Language Model Safety Under Parental Anxiety and Pressure in Pediatric Consultations

2025-12-17 · Vahideh Zolfaghari arxiv

Large language models (LLMs) are increasingly consulted by parents for pediatric guidance, yet their safety under real-world adversarial pressures is poorly understood. Anxious parents often use urgent language that can …

On Safety Risks in Experience-Driven Self-Evolving Agents

2026-04-18 · Weixiang Zhao, Yichen Zhang, Yingshuo Wang, Yang Deng 외 arxiv

Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduces underexplored safety risks. In this s…

How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity

2025-11-11 · Zihan Ma, Dongsheng Zhu, Shudong Liu, Taolin Zhang 외 arxiv

Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within complex tasks. We address this gap with a t…

Parallax: Why AI Agents That Think Must Never Act

2026-04-14 · Joel Fokou arxiv

Autonomous AI agents are rapidly transitioning from experimental tools to operational infrastructure, with projections that 80% of enterprise applications will embed AI copilots by the end of 2026. As agents gain the abi…

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

2026-05-26 · Aman Priyanshu, Supriti Vijay, Esha Pahwa arxiv

LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongside other agents. We introduce a Moltbook-style simulation platform w…