paper-with-me

홈 › Papers

A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models

2026-04-06 · Tianmeng Fang, Yong Wang, Zetai Kong, Zengzhen Su, Jun Wang, Chengjin Yu, Wei Wang arxiv

Multimodal large language models have become an important infrastructure for unified processing of visual and linguistic tasks. However, such models are highly susceptible to backdoor implantation during supervised fine-tuning and will steadily output the attacker's predefined harmful responses once a specific trigger pattern is activated. The core challenge of backdoor defense lies in suppressing attack success under low poisoning ratios while preserving the model's normal generation ability. These two objectives are inherently conflicting. Strong suppression often degrades benign performance, whereas weak regularization fails to mitigate backdoor behaviors. To this end, we propose a unified defense framework based on patch augmentation and cross-view regularity, which simultaneously constrains the model's anomalous behaviors in response to triggered patterns from both the feature representation and output distribution levels. Specifically, patch-level data augmentation is combined with cross-view output difference regularization to exploit the fact that backdoor responses are abnormally invariant to non-semantic perturbations and to proactively pull apart the output distributions of the original and perturbed views, thereby significantly suppressing the success rate of backdoor triggering. At the same time, we avoid over-suppression of the model during defense by imposing output entropy constraints, ensuring the quality of normal command generation. Experimental results across three models, two tasks, and six attacks show that our proposed defense method effectively reduces the attack success rate while maintaining a high level of normal text generation capability. Our work enables the secure, controlled deployment of large-scale multimodal models in realistic low-frequency poisoning and covert triggering scenarios.

📄 PDF Abstract BibTeX arXiv:2604.04488

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationText Generation

Similar Papers 제목 키워드 기반

PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers

2026-04-21 · Dazhuang Liu, Yanqi Qiao, Rui Wang, Kaitai Liang 외 arxiv

Vision Transformers (ViTs) have achieved remarkable success across vision tasks, yet recent studies show they remain vulnerable to backdoor attacks. Existing patch-wise attacks typically assume a single fixed trigger loc…

Patcher: Post-Hoc Patching of Backdoored Large Language Models

2026-06-02 · Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu 외 arxiv

Large language models remain vulnerable to jailbreak backdoor attacks, where adversaries poison safety alignment data to embed hidden triggers that bypass safety mechanisms. Existing defenses often require comprehensive …

Stealthy Patch-Wise Backdoor Attack in 3D Point Cloud via Curvature Awareness

2025-03-12 · Yu Feng, Dingxin Zhang, Runkai Zhao, Yong Xia 외

Backdoor attacks pose a severe threat to deep neural networks (DNN) by implanting hidden backdoors that can be activated with predefined triggers to manipulate model behaviors maliciously. Existing 3D point cloud backdoo…

Backdoor Attack

PatchBackdoor: Backdoor Attack against Deep Neural Networks without Model Modification

2023-08-22 · Yizhen Yuan, Rui Kong, Shenghao Xie, Yuanchun Li 외

Backdoor attack is a major threat to deep learning systems in safety-critical scenarios, which aims to trigger misbehavior of neural network models under attacker-controlled conditions. However, most backdoor attacks hav…

Adversarial AttackBackdoor AttackReal-World Adversarial Attack

Exploring Robustness of Visual State Space model against Backdoor Attacks

2024-08-21 · Cheng-Yi Lee, Cheng-Chang Tsai, Chia-Mu Yu, Chun-Shien Lu

Visual State Space Model (VSS) has demonstrated remarkable performance in various computer vision tasks. However, in the process of development, backdoor attacks have brought severe challenges to security. Such attacks c…