paper-with-me

홈 › Papers

REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space

2024-06-13 · Tomer Ashuach, Martin Tutek, Yonatan Belinkov

Language models (LMs) risk inadvertently memorizing and divulging sensitive or personally identifiable information (PII) seen in training data, causing privacy concerns. Current approaches to address this issue involve costly dataset scrubbing, or model filtering through unlearning and model editing, which can be bypassed through extraction attacks. We propose REVS, a novel non-gradient-based method for unlearning sensitive information from LMs. REVS identifies and modifies a small subset of neurons relevant for constituent tokens which form sensitive information. To adequately evaluate our method on truly sensitive information, we curate two datasets: an email dataset naturally memorized by Llama-3-8B and GPT-J-6B, and a synthetic social security number dataset that we tune the models to memorize. Compared to other methods, REVS demonstrates superior performance in unlearning sensitive information and robustness to extraction attacks, while retaining underlying model integrity.

📄 PDF Abstract BibTeX arXiv:2406.09325

Code (0)

등록된 구현이 없습니다.

Tasks

Model Editing

Similar Papers 제목 키워드 기반

iShumei-Chinchunmei at SemEval-2025 Task 4: A balanced forgetting and retention multi-task framework using effective unlearning loss

2025-07-22 · Yujian Sun, Tian Li arxiv

As the Large Language Model (LLM) gains widespread adoption, increasing attention has been given to the challenge of making LLM forget non-compliant data memorized during its pre-training. Machine Unlearning focuses on e…

Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs

2024-08-13 · Sungmin Cha, Sungjun Cho, Dasol Hwang, Moontae Lee

Large Language Models (LLMs) have demonstrated strong reasoning and memorization capabilities via pretraining on massive textual corpora. However, this poses risk of privacy and copyright violations, highlighting the nee…

Machine UnlearningMemorizationText Generation

Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models

2024-02-08 · Lingzhi Wang, Xingshan Zeng, Jinsong Guo, Kam-Fai Wong 외

This paper explores Machine Unlearning (MU), an emerging field that is gaining increased attention due to concerns about neural models unintentionally remembering personal or sensitive information. We present SeUL, a nov…

Computational EfficiencyLanguage ModellingMachine UnlearningMemorization

AILS-NTUA at SemEval-2025 Task 4: Parameter-Efficient Unlearning for Large Language Models using Data Chunking

2025-03-04 · Iraklis Premptis, Maria Lymperaiou, Giorgos Filandrianos, Orfeas Menis Mastromichalakis 외

The Unlearning Sensitive Content from Large Language Models task aims to remove targeted datapoints from trained models while minimally affecting their general knowledge. In our work, we leverage parameter-efficient, gra…

ChunkingGeneral Knowledge

Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning

2026-01-22 · Xinjie Zhou, Zhihui Yang, Lechao Cheng, Sai Wu 외 arxiv

Large language models (LLMs) exhibit powerful capabilities but risk memorizing sensitive personally identifiable information (PII) from their training data, posing significant privacy concerns. While machine unlearning t…