paper-with-me

Papers

Relationship-Aware Safety Unlearning for Multimodal LLMs

2026-03-15 · Vishnu Narayanan Anilkumar, Abhijith Sreesylesh Babu, Trieu Hai Vo, Mohankrishna Kolla, Alexander Cuneo arxiv

Generative multimodal models can exhibit safety failures that are inherently relational: two benign concepts can become unsafe when linked by a specific action or relation (e.g., child-drinking-wine). Existing unlearning and concept-erasure approaches often target isolated concepts or image-text pairs, which can cause collateral damage to benign uses of the same objects and relations. We propose relationship-aware safety unlearning: a framework that explicitly represents unsafe object-relation-object (O-R-O) tuples and applies targeted parameter-efficient edits (LoRA) to suppress unsafe tuples while preserving object marginals and safe neighboring relations. We include CLIP-based experiments and robustness evaluation under paraphrase, contextual, and out-of-distribution image attacks.

📄 PDF Abstract BibTeX arXiv:2603.14185

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

2025-02-18 · Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan 외

As Multimodal Large Language Models (MLLMs) develop, their potential security issues have become increasingly prominent. Machine Unlearning (MU), as an effective strategy for forgetting specific knowledge in training dat…

Machine UnlearningVisual Question Answering (VQA)

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation

2025-05-01 · Vaidehi Patil, Yi-Lin Sung, Peter Hase, Jie Peng 외

LLMs trained on massive datasets may inadvertently acquire sensitive information such as personal details and potentially harmful content. This risk is further heightened in multimodal LLMs as they integrate information …

Question AnsweringSpecificityVisual Question AnsweringVisual Question Answering (VQA)

Hierarchy-Aware Multimodal Unlearning for Medical AI

2025-12-10 · Fengli Wu, Vaidehi Patil, Jaehong Yoon, Yue Zhang 외 arxiv

Pretrained Multimodal Large Language Models (MLLMs) are increasingly used in sensitive domains such as medical AI, where privacy regulations like HIPAA and GDPR require specific removal of individuals' or institutions' d…

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

2026-05-15 · Jiahui Guang, Haiyan Wang, Yingjie Zhu, Cuiyun Gao 외 arxiv

Multimodal large language models (MLLMs) may memorize sensitive cross-modal information during pretraining, making machine unlearning (MU) crucial. Existing methods typically evaluate unlearning effectiveness based on ou…

VLSBench: Unveiling Visual Leakage in Multimodal Safety

2024-11-29 · Xuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang 외

Safety concerns of Multimodal large language models (MLLMs) have gradually become an important problem in various applications. Surprisingly, previous works indicate a counterintuitive phenomenon that using textual unlea…