paper-with-me

홈 › Papers

STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents

2025-09-30 · Jing-Jing Li, Jianfeng He, Chao Shang, Devang Kulshreshtha, Xun Xian, Yi Zhang, Hang Su, Sandesh Swamy, Yanjun Qi arxiv

As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining (STAC), a novel multi-turn attack framework that exploits agent tool use. STAC chains together tool calls that each appear harmless in isolation but, when combined, collectively enable harmful operations that only become apparent at the final execution step. We apply our framework to automatically generate and systematically evaluate 483 STAC cases, featuring 1,352 sets of user-agent-environment interactions and spanning diverse domains, tasks, agent types, and 10 failure modes. Our evaluations show that state-of-the-art LLM agents, including GPT-4.1, are highly vulnerable to STAC, with attack success rates (ASR) exceeding 90% in most cases. The core design of STAC's automated framework is a closed-loop pipeline that synthesizes executable multi-step tool chains, validates them through in-environment execution, and reverse-engineers stealthy multi-turn prompts that reliably induce agents to execute the verified malicious sequence. We further perform defense analysis against STAC and find that existing prompt-based defenses provide limited protection. To address this gap, we propose a new reasoning-driven defense prompt that achieves far stronger protection, cutting ASR by up to 28.8%. These results highlight a crucial gap: defending tool-enabled agents requires reasoning over entire action sequences and their cumulative effects, rather than evaluating isolated prompts or responses.

📄 PDF Abstract BibTeX arXiv:2509.25624

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

INNOCENT FAVOUR PRAISE

2025-06-14 · 06/14 2025 6 · INNOCENT FAVOUR PRAISE

SHORT BIOGRAPHY My Name is INNOCENT FAVOUR PRAISE I I'M FROM IMO STATE I WAS BORN 11/10/2009, I was also Born In Ilorin Kwara State, Also Born In The Family Of Mr And Miss INNOCENT, I'M Introduced Into A Business Called…

Marketing

Poison Forensics: Traceback of Data Poisoning Attacks in Neural Networks

2021-10-13 · Shawn Shan, Arjun Nitin Bhagoji, Haitao Zheng, Ben Y. Zhao

In adversarial machine learning, new defenses against attacks on deep learning systems are routinely broken soon after their release by more powerful attacks. In this context, forensic tools can offer a valuable compleme…

Data PoisoningMalware Classification

Infrastructure-based End-to-End Learning and Prevention of Driver Failure

2023-03-21 · Noam Buckman, Shiva Sreeram, Mathias Lechner, Yutong Ban 외

Intelligent intersection managers can improve safety by detecting dangerous drivers or failure modes in autonomous vehicles, warning oncoming vehicles as they approach an intersection. In this work, we present FailureNet…

Autonomous Vehicles

Efficient Black-box Assessment of Autonomous Vehicle Safety

2019-12-08 · Justin Norden, Matthew O'Kelly, Aman Sinha

While autonomous vehicle (AV) technology has shown substantial progress, we still lack tools for rigorous and scalable testing. Real-world testing, the $\textit{de-facto}$ evaluation method, is dangerous to the public. M…

Whose Bias?

2021-11-19 · Vasudha Jain, Mark Whitmeyer

Law enforcement acquires costly evidence with the aim of securing the conviction of a defendant, who is convicted if a decision-maker's belief exceeds a certain threshold. Either law enforcement or the decision-maker is …