paper-with-me

Papers

Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

2025-08-25 · Qiming Guo, Jinwen Tang, Xingran Huang arxiv

We introduce Advertisement Embedding Attacks (AEA), a new class of LLM security threats that stealthily inject promotional or malicious content into model outputs and AI agents. AEA operate through two low-cost vectors: (1) hijacking third-party service-distribution platforms to prepend adversarial prompts, and (2) publishing back-doored open-source checkpoints fine-tuned with attacker data. Unlike conventional attacks that degrade accuracy, AEA subvert information integrity, causing models to return covert ads, propaganda, or hate speech while appearing normal. We detail the attack pipeline, map five stakeholder victim groups, and present an initial prompt-based self-inspection defense that mitigates these injections without additional model retraining. Our findings reveal an urgent, under-addressed gap in LLM security and call for coordinated detection, auditing, and policy responses from the AI-safety community.

📄 PDF Abstract BibTeX arXiv:2508.17674

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Attacking Vision-Language Computer Agents via Pop-ups

2024-11-04 · Yanzhe Zhang, Tao Yu, Diyi Yang

Autonomous agents powered by large vision and language models (VLM) have demonstrated significant potential in completing daily computer tasks, such as browsing the web to book travel and operating desktop software, whic…

Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

2024-02-14 · Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel 외

Current research in adversarial robustness of LLMs focuses on discrete input manipulations in the natural language space, which can be directly transferred to closed-source models. However, this approach neglects the ste…

Adversarial RobustnessSafety Alignment

The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks

2024-10-20 · Daniel Ayzenshteyn, Roy Weiss, Yisroel Mirsky

As large language models (LLMs) continue to evolve, their potential use in automating cyberattacks becomes increasingly likely. With capabilities such as reconnaissance, exploitation, and command execution, LLMs could so…

Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents

2024-02-17 · Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 외

Driven by the rapid development of Large Language Models (LLMs), LLM-based agents have been developed to handle various real-world applications, including finance, healthcare, and shopping, etc. It is crucial to ensure t…

Backdoor Attackbackdoor defenseData Poisoning

Information Leakage from Embedding in Large Language Models

2024-05-20 · Zhipeng Wan, Anda Cheng, Yinggui Wang, Lei Wang

The widespread adoption of large language models (LLMs) has raised concerns regarding data privacy. This study aims to investigate the potential for privacy invasion through input reconstruction attacks, in which a malic…