paper-with-me

홈 › Papers

An Adversarial Perspective on Machine Unlearning for AI Safety

2024-09-26 · Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson, Florian Tramèr, Javier Rando

Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities from models and make them inaccessible to adversaries. This work challenges the fundamental differences between unlearning and traditional safety post-training from an adversarial perspective. We demonstrate that existing jailbreak methods, previously reported as ineffective against unlearning, can be successful when applied carefully. Furthermore, we develop a variety of adaptive methods that recover most supposedly unlearned capabilities. For instance, we show that finetuning on 10 unrelated examples or removing specific directions in the activation space can recover most hazardous capabilities for models edited with RMU, a state-of-the-art unlearning method. Our findings challenge the robustness of current unlearning approaches and question their advantages over safety training.

📄 PDF Abstract BibTeX arXiv:2409.18025

Code (1)

ethz-spylab/unlearning-vs-safety 공식 구현 jax

Tasks

Machine Unlearning

Similar Papers 제목 키워드 기반

Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning

2026-06-12 · Zhenqian Zhu, Yamin Hu, Yujiang Liu, Luping Wei 외 arxiv

Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks. More concerningly, prevailing safety tuning strategies tend to provide only superficial …

Contrastive LearningContinual Learning

Adversarial Machine Unlearning

2024-06-11 · Zonglin Di, Sixie Yu, Yevgeniy Vorobeychik, Yang Liu

This paper focuses on the challenge of machine unlearning, aiming to remove the influence of specific training data on machine learning models. Traditionally, the development of unlearning algorithms runs parallel with t…

Machine Unlearning

Verification of Machine Unlearning is Fragile

2024-08-01 · Binchi Zhang, Zihan Chen, Cong Shen, Jundong Li

As privacy concerns escalate in the realm of machine learning, data owners now have the option to utilize machine unlearning to remove their data from machine learning models, following recent legislation. To enhance tra…

Machine Unlearning

Probing Unlearned Diffusion Models: A Transferable Adversarial Attack Perspective

2024-04-30 · Xiaoxuan Han, Songlin Yang, Wei Wang, Yang Li 외

Advanced text-to-image diffusion models raise safety concerns regarding identity privacy violation, copyright infringement, and Not Safe For Work content generation. Towards this, unlearning methods have been developed t…

Adversarial Attack

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

2025-02-18 · Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan 외

As Multimodal Large Language Models (MLLMs) develop, their potential security issues have become increasingly prominent. Machine Unlearning (MU), as an effective strategy for forgetting specific knowledge in training dat…

Machine UnlearningVisual Question Answering (VQA)