paper-with-me

Papers

CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models

2026-05-04 · Ji Guo, Xiaolong Qin, Cencen Liu, Jielei Wang, Jierun Chen, Wenbo Jiang arxiv

Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their applications become increasingly widespread, recent studies have revealed that VLMs are vulnerable to backdoor attacks. Existing backdoor attacks on VLMs primarily rely on data poisoning by adding visual triggers and modifying text labels, where the induced image-text mismatch makes poisoned samples easy to detect. To address this limitation, we propose the Clean-Label Backdoor Attack on VLMs via Diffusion Models (CBV), which leverages diffusion models to generate natural poisoned examples via score matching. Specifically, CBV modifies the score during the reverse generation process of the diffusion model to guide the generation of poisoned samples that contain triggered image features. To further enhance the effectiveness of the attack, we incorporate the textual information of the triggered images as multimodal guidance during generation. Moreover, to enhance stealthiness, we introduce a GradCAM-guided Mask (GM) that restricts modifications to only the most semantically important regions, rather than the entire image. We evaluate our method on MSCOCO and VQA v2 with four representative VLMs, achieving over 80% ASR while preserving normal functionality.

📄 PDF Abstract BibTeX arXiv:2605.02202

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringImage Captioning

Similar Papers 제목 키워드 기반

Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification

2025-08-21 · Onur Alp Kirci, M. Emre Gursoy arxiv

Backdoor attacks pose a significant threat to the integrity of text classification models used in natural language processing. While several dirty-label attacks that achieve high attack success rates (ASR) have been prop…

Text Classification

Kallima: A Clean-label Framework for Textual Backdoor Attacks

2022-06-03 · Xiaoyi Chen, Yinpeng Dong, Zeyu Sun, Shengfang Zhai 외

Although Deep Neural Network (DNN) has led to unprecedented progress in various natural language processing (NLP) tasks, research shows that deep models are extremely vulnerable to backdoor attacks. The existing backdoor…

Large Language Models Are Better Adversaries: Exploring Generative Clean-Label Backdoor Attacks Against Text Classifiers

2023-10-28 · Wencong You, Zayd Hammoudeh, Daniel Lowd

Backdoor attacks manipulate model predictions by inserting innocuous triggers into training and test data. We focus on more realistic and more challenging clean-label attacks where the adversarial training examples are c…

Megatron: Evasive Clean-Label Backdoor Attacks against Vision Transformer

2024-12-06 · Xueluan Gong, Bowei Tian, Meng Xue, Shuike Li 외

Vision transformers have achieved impressive performance in various vision-related tasks, but their vulnerability to backdoor attacks is under-explored. A handful of existing works focus on dirty-label attacks with wrong…

Backdoor Attack

Enhancing Clean Label Backdoor Attack with Two-phase Specific Triggers

2022-06-10 · Nan Luo, Yuanzhang Li, Yajie Wang, Shangbo Wu 외

Backdoor attacks threaten Deep Neural Networks (DNNs). Towards stealthiness, researchers propose clean-label backdoor attacks, which require the adversaries not to alter the labels of the poisoned training datasets. Clea…

Backdoor Attackbackdoor defenseVocal Bursts Valence Prediction