paper-with-me

Papers

CoDefend: Cross-Modal Collaborative Defense via Diffusion Purification and Prompt Optimization

2025-10-13 · Fengling Zhu, Boshi Liu, Jingyu Hua, Sheng Zhong arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success in tasks such as image captioning, visual question answering, and cross-modal reasoning by integrating visual and textual modalities. However, their multimodal nature also exposes them to adversarial threats, where attackers can perturb either modality or both jointly to induce harmful, misleading, or policy violating outputs. Existing defense strategies, such as adversarial training and input purification, face notable limitations: adversarial training typically improves robustness only against known attacks while incurring high computational costs, whereas conventional purification approaches often suffer from degraded image quality and insufficient generalization to complex multimodal tasks. In this work, we focus on defending the visual modality, which frequently serves as the primary entry point for adversarial manipulation. We propose a supervised diffusion based denoising framework that leverages paired adversarial clean image datasets to fine-tune diffusion models with directional, task specific guidance. Unlike prior unsupervised purification methods such as DiffPure, our approach achieves higher quality reconstructions while significantly improving defense robustness in multimodal tasks. Furthermore, we incorporate prompt optimization as a complementary defense mechanism, enhancing resistance against diverse and unseen attack strategies. Extensive experiments on image captioning and visual question answering demonstrate that our method not only substantially improves robustness but also exhibits strong transferability to unknown adversarial attacks. These results highlight the effectiveness of supervised diffusion based denoising for multimodal defense, paving the way for more reliable and secure deployment of MLLMs in real world applications.

📄 PDF Abstract BibTeX arXiv:2510.11096

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringImage Captioning

Similar Papers 제목 키워드 기반

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models

2026-06-25 · Abrar Alotaibi, Moataz Ahmed arxiv

Adversarial evaluation of AI systems has matured along four largely disconnected tracks: diffusion-based attacks on text and large language models (LLMs), diffusion-based attacks on image classifiers, jailbreak pipelines…

Collaborative Diffusion for Multi-Modal Face Generation and Editing

2023-04-20 · CVPR 2023 1 · Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei Liu

Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condit…

DenoisingFace Generation

Reasoning with Autoregressive-Diffusion Collaborative Thoughts

2026-02-02 · Mu Yuan, Liekang Zeng, Guoliang Xing, Lan Zhang 외 arxiv

Autoregressive and diffusion models represent two complementary generative paradigms. Autoregressive models excel at sequential planning and constraint composition, yet struggle with tasks that require explicit spatial o…

Question AnsweringSpatial Reasoning

When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models

2026-03-06 · Qitong Wang, Haoran Dai, Haotian Zhang, Christopher Rasmussen 외 arxiv

While diffusion models have revolutionized visual content generation, their rapid adoption has underscored the critical need to investigate vulnerabilities, e.g., to backdoor attacks. In multimodal diffusion models, it i…

Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation

2025-12-22 · Guoli Jia, Junyao Hu, Xinwei Long, Kai Tian 외 arxiv

Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emotion in advertising, emotion-oriented ima…

Image Generation