paper-with-me

Papers

DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent

2025-02-18 · Pengyu Zhu, Zhenhong Zhou, Yuanhe Zhang, Shilinlu Yan, Kun Wang, Sen Su

As LLM-based agents become increasingly prevalent, backdoors can be implanted into agents through user queries or environment feedback, raising critical concerns regarding safety vulnerabilities. However, backdoor attacks are typically detectable by safety audits that analyze the reasoning process of agents. To this end, we propose a novel backdoor implantation strategy called \textbf{Dynamically Encrypted Multi-Backdoor Implantation Attack}. Specifically, we introduce dynamic encryption, which maps the backdoor into benign content, effectively circumventing safety audits. To enhance stealthiness, we further decompose the backdoor into multiple sub-backdoor fragments. Based on these advancements, backdoors are allowed to bypass safety audits significantly. Additionally, we present AgentBackdoorEval, a dataset designed for the comprehensive evaluation of agent backdoor attacks. Experimental results across multiple datasets demonstrate that our method achieves an attack success rate nearing 100\% while maintaining a detection rate of 0\%, illustrating its effectiveness in evading safety audits. Our findings highlight the limitations of existing safety mechanisms in detecting advanced attacks, underscoring the urgent need for more robust defenses against backdoor threats. Code and data are available at https://github.com/whfeLingYu/DemonAgent.

📄 PDF Abstract BibTeX arXiv:2502.12575

Code (1)

whfelingyu/demonagent 공식 구현

Similar Papers 제목 키워드 기반

Enhancing Clean Label Backdoor Attack with Two-phase Specific Triggers

2022-06-10 · Nan Luo, Yuanzhang Li, Yajie Wang, Shangbo Wu 외

Backdoor attacks threaten Deep Neural Networks (DNNs). Towards stealthiness, researchers propose clean-label backdoor attacks, which require the adversaries not to alter the labels of the poisoned training datasets. Clea…

Backdoor Attackbackdoor defenseVocal Bursts Valence Prediction

Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness

2025-10-09 · Jiyang Qiu, Xinbei Ma, Yunqing Xu, Zhuosheng Zhang 외 arxiv

The rapid deployment of large language model (LLM)-based agents in real-world applications has raised serious concerns about their trustworthiness. In this work, we reveal the security and robustness vulnerabilities of t…

SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks

2024-05-19 · Xuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein 외

Modern NLP models are often trained on public datasets drawn from diverse sources, rendering them vulnerable to data poisoning attacks. These attacks can manipulate the model's behavior in ways engineered by the attacker…

Data Poisoning

Lightweight and Fast Backdoor Model Detection

2026-05-17 · Yinbo Yu, Jing Fang, Xuewen Zhang, Chunwei Tian 외 arxiv

Deep neural networks (DNN), despite their remarkable performance, are highly vulnerable to backdoor attacks. Existing defenses mainly rely on activation anomaly analysis or trigger reverse engineering and often require c…

Cooperative Decentralized Backdoor Attacks on Vertical Federated Learning

2025-01-16 · Seohyun Lee, Wenzhi Fang, Anindya Bijoy Das, Seyyedali Hosseinalipour 외

Federated learning (FL) is vulnerable to backdoor attacks, where adversaries alter model behavior on target classification labels by embedding triggers into data samples. While these attacks have received considerable at…

Backdoor AttackFederated LearningMetric LearningVertical Federated Learning