paper-with-me

홈 › Papers

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

2026-07-08 · Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil arxiv

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retraining after deletion requests or policy updates is often impractical, and targeted forgetting remains difficult because knowledge is distributed across shared representations. Multimodal unlearning addresses this challenge by enabling selective removal across modalities while retaining overall utility. This survey offers a unified, system-oriented view of multimodal unlearning across vision, language, audio, and video, grounded in recent advances, emerging applications, and open problems. Our taxonomy enables systematic comparison across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. This survey highlights open problems and practical considerations to support future research and deployment of multimodal unlearning. We release a curated repository: https://smsnobin77.github.io/Awesome-Multimodal-Unlearning/

📄 PDF Abstract BibTeX arXiv:2607.07907

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SineProject: Machine Unlearning for Stable Vision Language Alignment

2025-11-23 · Arpit Garg, Hemanth Saratchandran, Simon Lucey arxiv

Multimodal Large Language Models (MLLMs) increasingly need to forget specific knowledge such as unsafe or private information without requiring full retraining. However, existing unlearning methods often disrupt vision l…

On the Robustness of Machine Unlearning for Vision-Language Models

2026-05-26 · Yujie Lin, Kaidi Jia, Jiayao Ma, Chengyi Yang 외 arxiv

Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systematic survey and robustness analysis of VL…

Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models

2026-03-23 · Hyundong Jin, Dongyoon Han, Eunwoo Kim arxiv

Continual unlearning poses the challenge of enabling large vision-language models to selectively refuse specific image-instruction pairs in response to sequential deletion requests, while preserving general utility. Howe…

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

2026-05-15 · Jiahui Guang, Haiyan Wang, Yingjie Zhu, Cuiyun Gao 외 arxiv

Multimodal large language models (MLLMs) may memorize sensitive cross-modal information during pretraining, making machine unlearning (MU) crucial. Existing methods typically evaluate unlearning effectiveness based on ou…

Hierarchy-Aware Multimodal Unlearning for Medical AI

2025-12-10 · Fengli Wu, Vaidehi Patil, Jaehong Yoon, Yue Zhang 외 arxiv

Pretrained Multimodal Large Language Models (MLLMs) are increasingly used in sensitive domains such as medical AI, where privacy regulations like HIPAA and GDPR require specific removal of individuals' or institutions' d…