paper-with-me

홈 › Papers

Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models

2024-07-02 · Xu Han, Linghao Jin, Xuezhe Ma, Xiaofeng Liu

Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are restricted by patient privacy concerns, potentially containing noise that can adversely affect downstream performance. Moreover, the growing reliance on multi-modal generation exacerbates this issue because of its susceptibility to adversarial attacks. To investigate how VLMs trained on adversarial noisy data perform on downstream medical tasks, we first craft noisy upstream datasets using multi-modal adversarial attacks. Through our comprehensive analysis, we unveil that moderate noise enhances model robustness and transferability, but increasing noise levels negatively impact downstream task performance. To mitigate this issue, we propose rectify adversarial noise (RAN) framework, a recipe designed to effectively defend adversarial attacks and rectify the influence of upstream noise during fine-tuning.

📄 PDF Abstract BibTeX arXiv:2407.02716

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs

2026-02-21 · Chun Yan Ryan Kan, Tommy Tran, Vedant Yadav, Ava Cai 외 arxiv

Defending LLMs against adversarial jailbreak attacks remains an open challenge. Existing defenses rely on binary classifiers that fail when adversarial input falls outside the learned decision boundary, and repeated fine…

Density Estimation

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

2026-06-06 · Haoming Wen, Shi Chen, Qingyu Shi, Siyuan Liu 외 arxiv

Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised finetuning (SFT) on poisoned datasets. E…

LiBRe: A Practical Bayesian Approach to Adversarial Detection

2021-03-27 · CVPR 2021 1 · Zhijie Deng, Xiao Yang, Shizhen Xu, Hang Su 외

Despite their appealing flexibility, deep neural networks (DNNs) are vulnerable against adversarial examples. Various adversarial defense strategies have been proposed to resolve this problem, but they typically demonstr…

Adversarial DefenseUncertainty Quantification

DLADiff: A Dual-Layer Defense Framework against Fine-Tuning and Zero-Shot Customization of Diffusion Models

2025-11-25 · Jun Jia, Hongyi Miao, Yingjie Zhou, Linhan Cao 외 arxiv

With the rapid advancement of diffusion models, a variety of fine-tuning methods have been developed, enabling high-fidelity image generation with high similarity to the target content using only 3 to 5 training images. …

Image Generation

Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs

2024-06-07 · Fan Liu, Zhao Xu, Hao liu

Although safely enhanced Large Language Models (LLMs) have achieved remarkable success in tackling various complex tasks in a zero-shot manner, they remain susceptible to jailbreak attacks, particularly the unknown jailb…

Prompt Learning