paper-with-me

홈 › Papers

OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models

2025-05-07 · Xiaoyu Xu, Minxin Du, Qingqing Ye, Haibo Hu

Large language models (LLMs) trained over extensive corpora risk memorizing sensitive, copyrighted, or toxic content. To address this, we propose OBLIVIATE, a robust unlearning framework that removes targeted data while preserving model utility. The framework follows a structured process: extracting target tokens, building retain sets, and fine-tuning with a tailored loss function comprising three components -- masking, distillation, and world fact. Using low-rank adapters (LoRA), it ensures efficiency without compromising unlearning quality. We conduct experiments on multiple datasets, including the Harry Potter series, WMDP, and TOFU, using a comprehensive suite of metrics: forget quality (new document-level memorization score), model utility, and fluency. Results demonstrate its effectiveness in resisting membership inference attacks, minimizing the impact on retained data, and maintaining robustness across diverse scenarios.

📄 PDF Abstract BibTeX arXiv:2505.04416

Code (0)

등록된 구현이 없습니다.

Tasks

Machine UnlearningMemorization

Similar Papers 제목 키워드 기반

Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models

2025-02-20 · Mark Russinovich, Ahmed Salem

Recent copyright agreements between AI companies and content creators underscore the need for fine-grained control over language models' ability to reproduce copyrighted text. Existing defenses-ranging from aggressive un…

HellaSwagMemorizationMMLUTruthfulQA+1

DeepObliviate: A Powerful Charm for Erasing Data Residual Memory in Deep Neural Networks

2021-05-13 · Yingzhe He, Guozhu Meng, Kai Chen, Jinwen He 외

Machine unlearning has great significance in guaranteeing model security and protecting user privacy. Additionally, many legal provisions clearly stipulate that users have the right to demand model providers to delete th…

Machine Unlearning

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

2025-08-06 · Yifan Li, Kun Zhou, Wayne Xin Zhao, Lei Fang 외 arxiv

As scaling up training data has significantly improved the general multimodal capabilities of Large Vision-Language Models (LVLMs), they still suffer from the hallucination issue, generating text that is inconsistent wit…

Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm

2024-09-21 · Jaehan Kim, Minkyoo Song, Seung Ho Na, Seungwon Shin

Parameter-efficient fine-tuning (PEFT) has become a key training strategy for large language models. However, its reliance on fewer trainable parameters poses security risks, such as task-agnostic backdoors. Despite thei…

backdoor defenseparameter-efficient fine-tuning

The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework

2025-10-29 · Aakriti Shah, Thai Le arxiv

Unlearning in large language models (LLMs) is crucial for managing sensitive data and correcting misinformation, yet evaluating its effectiveness remains an open problem. We investigate whether persuasive prompting can r…