paper-with-me

홈 › Papers

Identifying the Risks of LM Agents with an LM-Emulated Sandbox

2023-09-25 · Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, Tatsunori Hashimoto

Recent advances in Language Model (LM) agents and tool use, exemplified by applications like ChatGPT Plugins, enable a rich set of capabilities but also amplify potential risks - such as leaking private data or causing financial losses. Identifying these risks is labor-intensive, necessitating implementing the tools, setting up the environment for each test scenario manually, and finding risky cases. As tools and agents become more complex, the high cost of testing these agents will make it increasingly difficult to find high-stakes, long-tailed risks. To address these challenges, we introduce ToolEmu: a framework that uses an LM to emulate tool execution and enables the testing of LM agents against a diverse range of tools and scenarios, without manual instantiation. Alongside the emulator, we develop an LM-based automatic safety evaluator that examines agent failures and quantifies associated risks. We test both the tool emulator and evaluator through human evaluation and find that 68.8% of failures identified with ToolEmu would be valid real-world agent failures. Using our curated initial benchmark consisting of 36 high-stakes tools and 144 test cases, we provide a quantitative risk analysis of current LM agents and identify numerous failures with potentially severe outcomes. Notably, even the safest LM agent exhibits such failures 23.9% of the time according to our evaluator, underscoring the need to develop safer LM agents for real-world deployment.

📄 PDF Abstract BibTeX arXiv:2309.15817

Code (1)

ryoungj/toolemu 공식 구현

Tasks

Language Modellingvalid

Similar Papers 제목 키워드 기반

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

2026-03-01 · Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis 외 arxiv

Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, creating novel security risks. To mitigate these risks, agents are commonly depl…

HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions

2024-09-24 · Xuhui Zhou, Hyunwoo Kim, Faeze Brahman, Liwei Jiang 외

AI agents are increasingly autonomous in their interactions with human users and tools, leading to increased interactional safety risks. We present HAICOSYSTEM, a framework examining AI agent safety within diverse and co…

AI AgentNavigate

Fault-Tolerant Sandboxing for AI Coding Agents: A Transactional Approach to Safe Autonomous Execution

2025-12-14 · Boyang Yan arxiv

The transition of Large Language Models (LLMs) from passive code generators to autonomous agents introduces significant safety risks, specifically regarding destructive commands and inconsistent system states. Existing c…

Agent Safety Alignment via Reinforcement Learning

2025-07-11 · Zeyang Sha, Hanling Tian, Zhuoer Xu, Shiwen Cui 외 arxiv

The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents, empowered to execute external functions,…

Reinforcement Learning

LLM Agents Should Employ Security Principles

2025-05-29 · Kaiyuan Zhang, Zian Su, Pin-Yu Chen, Elisa Bertino 외

Large Language Model (LLM) agents show considerable promise for automating complex tasks using contextual reasoning; however, interactions involving multiple agents and the system's susceptibility to prompt injection and…

Large Language Model