paper-with-me

홈 › Papers

Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs

2024-06-17 · Swanand Ravindra Kadhe, Farhan Ahmed, Dennis Wei, Nathalie Baracaldo, Inkit Padhi

Large language models (LLMs) have shown to pose social and ethical risks such as generating toxic language or facilitating malicious use of hazardous knowledge. Machine unlearning is a promising approach to improve LLM safety by directly removing harmful behaviors and knowledge. In this paper, we propose "SPlit, UNlearn, MerGE" (SPUNGE), a framework that can be used with any unlearning method to amplify its effectiveness. SPUNGE leverages data attributes during unlearning by splitting unlearning data into subsets based on specific attribute values, unlearning each subset separately, and merging the unlearned models. We empirically demonstrate that SPUNGE significantly improves the performance of two recent unlearning methods on state-of-the-art LLMs while maintaining their general capabilities on standard academic benchmarks.

📄 PDF Abstract BibTeX arXiv:2406.11780

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeMachine Unlearning

Similar Papers 제목 키워드 기반

FairGU: Fairness-aware Graph Unlearning in Social Networks

2026-01-14 · Renqiang Luo, Yongshuai Yang, Huafei Huang, Qing Qing 외 arxiv

Graph unlearning has emerged as a critical mechanism for supporting sustainable and privacy-preserving social networks, enabling models to remove the influence of deleted nodes and thereby better safeguard user informati…

Graph Learning

When unlearning is free: leveraging low influence points to reduce computational costs

2025-12-04 · Anat Kleiman, Robert Fisher, Ben Deaner, Udi Wieder arxiv

As concerns around data privacy in machine learning grow, the ability to unlearn, or remove, specific data points from trained models becomes increasingly important. While state of the art unlearning methods have emerged…

Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures

2026-02-03 · Sangyeon Yoon, Hyesoo Hong, Wonje Jeung, Albert No arxiv

Machine unlearning aims to remove specific content from trained models while preserving overall performance. However, the phenomenon of benign relearning, in which forgotten information reemerges even from benign fine-tu…

Efficient Attribute Unlearning: Towards Selective Removal of Input Attributes from Feature Representations

2022-02-27 · Tao Guo, Song Guo, Jiewei Zhang, Wenchao Xu 외

Recently, the enactment of privacy regulations has promoted the rise of the machine unlearning paradigm. Existing studies of machine unlearning mainly focus on sample-wise unlearning, such that a learnt model will not ex…

AttributeFace RecognitionFairnessMachine Unlearning

Consistency-Aware Editing for Entity-level Unlearning in Language Models

2025-12-19 · Xiaoqi Han, Víctor Gutiérrez-Basulto, Ru Li, Xiaoli Li 외 arxiv

Large language models (LLMs) risk retaining sensitive, copyrighted, or harmful information from their training data. Entity-level unlearning addresses this issue by removing all knowledge of a specific entity while prese…