paper-with-me

Papers

How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity

2025-11-11 · Zihan Ma, Dongsheng Zhu, Shudong Liu, Taolin Zhang, Junnan Liu, Qingqiu Li, Minnan Luo, Songyang Zhang, Kai Chen arxiv

Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within complex tasks. We address this gap with a two-dimensional analysis of agent safety brittleness under the orthogonal pressures of intent concealment and task complexity. To enable this, we introduce OASIS (Orthogonal Agent Safety Inquiry Suite), a hierarchical benchmark with fine-grained annotations and a high-fidelity simulation sandbox. Our findings reveal two critical phenomena: safety alignment degrades sharply and predictably as intent becomes obscured, and a "Complexity Paradox" emerges, where agents seem safer on harder tasks only due to capability limitations. By releasing OASIS and its simulation environment, we provide a principled foundation for probing and strengthening agent safety in these overlooked dimensions.

📄 PDF Abstract BibTeX arXiv:2511.08487

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAEBE: Multi-Agent Emergent Behavior Framework

2025-06-03 · Sinem Erisken, Timothy Gothard, Martin Leitgab, Ram Potham

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEB…

CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation

2026-04-10 · Yushi Feng, Junye Du, Qifan Wang, Zizhan Ma 외 arxiv

Graphical user interface (GUI) agents powered by vision language models (VLMs) are rapidly moving from passive assistance to autonomous operation. However, this unrestricted action space exposes users to severe and irrev…

Multimodal ReasoningPrompt Engineering

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

2026-05-08 · Zhengyang Tang, Yi Zhang, Chenxin Li, Xin Lai 외 arxiv

When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may be avoided because the agent recognized the risk and chose the safe act…

Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment

2026-02-03 · Jingnan Zheng, Yanzhen Luo, Jingjun Xu, Bingnan Liu 외 arxiv

Large Language Models (LLMs) are increasingly deployed as agents that operate in real-world environments, introducing safety risks beyond linguistic harm. Existing agent safety evaluations rely on risk-oriented tasks tai…

A Safety and Security Framework for Real-World Agentic Systems

2025-11-27 · Shaona Ghosh, Barnaby Simkin, Kyriacos Shiarlis, Soumili Nandi 외 arxiv

This paper introduces a dynamic and actionable framework for securing agentic AI systems in enterprise deployment. We contend that safety and security are not merely fixed attributes of individual models but also emergen…

Red Teaming