paper-with-me

홈 › Papers

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

2026-04-07 · Peigui Qi, Kunsheng Tang, Yanpu Yu, Jialin Wu, Yide Song, Wenbo Zhou, Zhicong Huang, Cheng Hong, Weiming Zhang, Nenghai Yu arxiv

Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration. Existing defenses suffer from efficiency and robustness. To address these challenges, we first propose the Multimodal Aggregated Feature Extraction (MAFE) framework that enables CLIP to handle long text and fuse multimodal information into unified representations. Through empirical analysis of MAFE-extracted features, we discover distinct distributional patterns between benign and malicious prompts. Building upon this finding, we develop VLMShield, a lightweight safety detector that efficiently identifies multimodal malicious attacks as a plug-and-play solution. Extensive experiments demonstrate superior performance across multiple dimensions, including robustness, efficiency, and utility. Through our work, we hope to pave the way for more secure multimodal AI deployment. Code is available at this https URL.

📄 PDF Abstract BibTeX arXiv:2604.06502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

2026-02-16 · In Chong Choi, Jiacheng Zhang, Feng Liu, Yiliao Song arxiv

Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced to bypass safety alignment. However, effectively extending such attacks t…

DE-FIVE: Detecting Malicious Image Prompts via Fourier Features and Image Vector Embeddings

2026-06-22 · Xingwei Zhong, Varun Sharma, Kar Wai Fok, Vrizlynn L. L. Thing arxiv

Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual modalities expands the attack surface of VLMs, making them more suscepti…

MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks

2024-02-29 · Fangyuan Zhang, Huichi Zhou, Shuangjiao Li, Hongtao Wang

Deep neural networks have been proven to be vulnerable to adversarial examples and various methods have been proposed to defend against adversarial attacks for natural language processing tasks. However, previous defense…

SDD: Self-Degraded Defense against Malicious Fine-tuning

2025-07-27 · Zixuan Chen, Weikai Lu, Xin Lin, Ziqian Zeng arxiv

Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass thes…

Towards Transferable Defense Against Malicious Image Edits

2025-12-16 · Jie Zhang, Shuai Dong, Shiguang Shan, Xilin Chen arxiv

Recent approaches employing imperceptible perturbations in input images have demonstrated promising potential to counter malicious manipulations in diffusion-based image editing systems. However, existing methods suffer …

Image Editing