paper-with-me

홈 › Papers

Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs

2024-08-13 · Sungmin Cha, Sungjun Cho, Dasol Hwang, Moontae Lee

Large Language Models (LLMs) have demonstrated strong reasoning and memorization capabilities via pretraining on massive textual corpora. However, this poses risk of privacy and copyright violations, highlighting the need for efficient machine unlearning methods that remove sensitive data without retraining from scratch. While Gradient Ascent (GA) is commonly used to unlearn by reducing the likelihood of generating unwanted content, it leads to unstable optimization and catastrophic forgetting of retrained knowledge. We find that combining GA with low-rank adaptation results in poor trade-offs between computational cost and generative performance. To address these challenges, we propose Low-rank Knowledge Unlearning (LoKU), a novel framework that enables robust and efficient unlearning for LLMs. First, we introduce Inverted Hinge Loss, which suppresses unwanted tokens while maintaining fluency by boosting the probability of the next most likely token. Second, we develop a data-adaptive initialization for LoRA adapters via low-rank approximation weighted with relative Fisher information, thereby focusing updates on parameters critical for removing targeted knowledge. Experiments on the Training Data Extraction Challenge dataset using GPT-Neo models as well as on the TOFU benchmark with Phi-1.5B and Llama2-7B models demonstrate that our approach effectively removes sensitive information while maintaining reasoning and generative capabilities with minimal impact. Our implementation can be found in https://github.com/csm9493/efficient-llm-unlearning.

📄 PDF Abstract BibTeX arXiv:2408.06621

Code (1)

csm9493/efficient-llm-unlearning 공식 구현 pytorch

Tasks

Machine UnlearningMemorizationText Generation

Methods 이 논문이 사용한 방법론

Tofu 설명 없음
GA Genetic Algorithms are search algorithms that mimic Darwinian biological evolution in order to select and propagate better solutions.
GPT-Neo An implementation of model & data parallel GPT3-like models using the mesh-tensorflow…
Adapter 설명 없음

Similar Papers 제목 키워드 기반

Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges

2023-11-27 · Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang 외

In recent years, large language models (LLMs) have spurred a new research paradigm in natural language processing. Despite their excellent capability in knowledge-based question answering and reasoning, their potential t…

In-Context LearningMachine UnlearningQuestion AnsweringSurvey

ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMs

2026-03-02 · Xunlei Chen, Jinyu Guo, Yuang Li, Zhaokun Wang 외 arxiv

Large language models (LLMs) have advanced to encompass extensive knowledge across diverse domains. Yet controlling what a LLMs should not know is important for ensuring alignment and thus safe use. However, effective un…

To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models

2024-07-02 · Bozhong Tian, Xiaozhuan Liang, Siyuan Cheng, Qingbin Liu 외

Large Language Models (LLMs) trained on extensive corpora inevitably retain sensitive data, such as personal privacy information and copyrighted material. Recent advancements in knowledge unlearning involve updating LLM …

General Knowledge

UOE: Unlearning One Expert Is Enough For Mixture-of-experts LLMS

2024-11-27 · Haomin Zhuang, Yihua Zhang, Kehan Guo, Jinghan Jia 외

Recent advancements in large language model (LLM) unlearning have shown remarkable success in removing unwanted data-model influences while preserving the model's utility for legitimate knowledge. However, despite these …

Large Language ModelMixture-of-Experts

Intrinsic Evaluation of Unlearning Using Parametric Knowledge Traces

2024-06-17 · Yihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel 외

The task of "unlearning" certain concepts in large language models (LLMs) has attracted immense attention recently, due to its importance in mitigating undesirable model behaviours, such as the generation of harmful, pri…