paper-with-me

홈 › Papers

Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models

2026-08-02 · Junkai Lin, Junkai Chen, Siqi Hou, Yuhao He, Ruiqi Liu, Chenhan Jin, Shengze Xu, Tieyong Zeng arxiv

Multimodal large language models (MLLMs) exhibit strong vision--language capabilities but may also memorize and disclose sensitive information. Machine unlearning seeks to remove designated knowledge without retraining from scratch while preserving general utility. Existing privacy-oriented benchmarks primarily adopt profile-level deletion, whereas practical requests are often finer grained: a model should forget a specified attribute while retaining non-sensitive information about the same identity. We therefore introduce attribute-level MLLM unlearning as a finer-grained task and construct a benchmark spanning long-text, numeric, and short-text targets, multiple forget ratios, and diverse question types. Our evaluation reveals that target and retained attributes share identity-specific and visual evidence, making selective forgetting susceptible to residual leakage or collateral degradation; accordingly, existing methods exhibit unstable forgetting--retention trade-offs in this setting. To address this challenge, we propose Causal Localization and Retain-Aware Projection (CLRP), a lightweight training-free framework. CLRP uses activation patching to identify the layer that causally mediates target-attribute disclosure, then applies a retain-aware projection that removes the target-attribute subspace while preserving same-identity evidence. Experiments across multiple widely used MLLMs with distinct architectures and parameter scales demonstrate the effectiveness of CLRP.

📄 PDF Abstract BibTeX arXiv:2608.01008

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SALMUBench: A Benchmark for Sensitive Association-Level Multimodal Unlearning

2026-03-27 · Cai Selvas-Sala, Lei Kang, Lluis Gomez arxiv

As multimodal models like CLIP become integral to downstream systems, the need to remove sensitive information is critical. However, machine unlearning for contrastively-trained encoders remains underexplored, and existi…

Hierarchy-Aware Multimodal Unlearning for Medical AI

2025-12-10 · Fengli Wu, Vaidehi Patil, Jaehong Yoon, Yue Zhang 외 arxiv

Pretrained Multimodal Large Language Models (MLLMs) are increasingly used in sensitive domains such as medical AI, where privacy regulations like HIPAA and GDPR require specific removal of individuals' or institutions' d…

Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models

2026-08-31 · Kangwook Ko, Jaehyuk Jang, Wonjun Lee, Hee-Seon Kim 외 arxiv

Removing a specific individual's information from multimodal large language models (MLLMs) is often needed after deployment, but existing methods rely on a retain set, which is hardest to obtain at that point, and rebuil…

ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition

2026-05-14 · Shen Lin, Jing Lin, Junhao Dong, Piotr Koniusz 외 arxiv

Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove target knowledge without affecting unrelated semantics. This issue is esp…

Targeted Forgetting of Image Subgroups in CLIP Models

2025-01-01 · CVPR 2025 1 · Zeliang Zhang, Gaowen Liu, Charles Fleming, Ramana Rao Kompella 외

Foundation models (FMs) such as CLIP have demonstrated impressive zero-shot performance across various tasks by leveraging large-scale, unsupervised pre-training. However, they often inherit harmful or unwanted knowl…

Knowledge DistillationUnsupervised Pre-training