paper-with-me

홈 › Papers

Eight Methods to Evaluate Robust Unlearning in LLMs

2024-02-26 · Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, Dylan Hadfield-Menell

Machine unlearning can be useful for removing harmful capabilities and memorized text from large language models (LLMs), but there are not yet standardized methods for rigorously evaluating it. In this paper, we first survey techniques and limitations of existing unlearning evaluations. Second, we apply a comprehensive set of tests for the robustness and competitiveness of unlearning in the "Who's Harry Potter" (WHP) model from Eldan and Russinovich (2023). While WHP's unlearning generalizes well when evaluated with the "Familiarity" metric from Eldan and Russinovich, we find i) higher-than-baseline amounts of knowledge can reliably be extracted, ii) WHP performs on par with the original model on Harry Potter Q&A tasks, iii) it represents latent knowledge comparably to the original model, and iv) there is collateral unlearning in related domains. Overall, our results highlight the importance of comprehensive unlearning evaluation that avoids ad-hoc metrics.

📄 PDF Abstract BibTeX arXiv:2402.16835

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Unlearning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ALU: Agentic LLM Unlearning

2025-02-01 · Debdeep Sanyal, Murari Mandal

Information removal or suppression in large language models (LLMs) is a desired functionality, useful in AI regulation, legal compliance, safety, and privacy. LLM unlearning methods aim to remove information on demand fr…

Effective Skill Unlearning through Intervention and Abstention

2025-03-27 · Yongce Li, Chung-En Sun, Tsui-Wei Weng

Large language Models (LLMs) have demonstrated remarkable skills across various domains. Understanding the mechanisms behind their abilities and implementing controls over them is becoming increasingly important for deve…

General KnowledgeMathMMLU

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation

2025-05-01 · Vaidehi Patil, Yi-Lin Sung, Peter Hase, Jie Peng 외

LLMs trained on massive datasets may inadvertently acquire sensitive information such as personal details and potentially harmful content. This risk is further heightened in multimodal LLMs as they integrate information …

Question AnsweringSpecificityVisual Question AnsweringVisual Question Answering (VQA)

WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models

2024-10-23 · Jinghan Jia, Jiancheng Liu, Yihua Zhang, Parikshit Ram 외

The need for effective unlearning mechanisms in large language models (LLMs) is increasingly urgent, driven by the necessity to adhere to data regulations and foster ethical generative AI practices. Despite growing inter…

DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning

2025-11-08 · Yaxuan Wang, Chris Yuhao Liu, Quan Liu, Jinglong Pang 외 arxiv

Unlearning in Large Language Models (LLMs) is crucial for protecting private data and removing harmful knowledge. Most existing approaches rely on fine-tuning to balance unlearning efficiency with general language capabi…