paper-with-me

홈 › Papers

UnPII: Unlearning Personally Identifiable Information with Quantifiable Exposure Risk

2026-01-05 · Intae Jeon, Yujeong Kwon, Hyungjoon Koo arxiv

The ever-increasing adoption of Large Language Models in critical sectors like finance, healthcare, and government raises privacy concerns regarding the handling of sensitive Personally Identifiable Information (PII) during training. In response, regulations such as European Union's General Data Protection Regulation (GDPR) mandate the deletion of PII upon requests, underscoring the need for reliable and cost-effective data removal solutions. Machine unlearning has emerged as a promising direction for selectively forgetting data points. However, existing unlearning techniques typically apply a uniform forgetting strategy that neither accounts for the varying privacy risks posed by different PII attributes nor reflects associated business risks. In this work, we propose UnPII, the first PII-centric unlearning approach that prioritizes forgetting based on the risk of individual or combined PII attributes. To this end, we introduce the PII risk index (PRI), a composite metric that incorporates multiple dimensions of risk factors: identifiability, sensitivity, usability, linkability, permanency, exposability, and compliancy. The PRI enables a nuanced evaluation of privacy risks associated with PII exposures and can be tailored to align with organizational privacy policies. To support realistic assessment, we systematically construct a synthetic PII dataset (e.g., 1,700 PII instances) that simulates realistic exposure scenarios. UnPII seamlessly integrates with established unlearning algorithms, such as Gradient Ascent, Negative Preference Optimization, and Direct Preference Optimization, without modifying their underlying principles. Our experimental results demonstrate that UnPII achieves the improvements of accuracy up to 11.8%, utility up to 6.3%, and generalizability up to 12.4%, respectively, while incurring a modest fine-tuning overhead of 27.5% on average during unlearning.

📄 PDF Abstract BibTeX arXiv:2601.01786

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SemEval-2025 Task 4: Unlearning sensitive content from Large Language Models

2025-04-02 · Anil Ramakrishna, Yixin Wan, Xiaomeng Jin, Kai-Wei Chang 외

We introduce SemEval-2025 Task 4: unlearning sensitive content from Large Language Models (LLMs). The task features 3 subtasks for LLM unlearning spanning different use cases: (1) unlearn long form synthetic creative doc…

Form

REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space

2024-06-13 · Tomer Ashuach, Martin Tutek, Yonatan Belinkov

Language models (LMs) risk inadvertently memorizing and divulging sensitive or personally identifiable information (PII) seen in training data, causing privacy concerns. Current approaches to address this issue involve c…

Model Editing

Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation

2026-05-14 · Kyomin Hwang, Hyeonjin Kim, Sangyeon Cho, Nojun Kwak arxiv

While LLMs are increasingly used in commercial services, they pose privacy risks such as leakage of sensitive personally identifiable information (PII). For LLMs trained on multilingual corpora, Multilingual Machine Unle…

Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs

2026-01-07 · Dinesh Srivasthav P, Ashok Urlana, Rahul Mishra, Bala Mallikarjunarao Garlapati 외 arxiv

Machine unlearning aims to selectively remove the influence of specific training samples to satisfy privacy regulations such as the GDPR's 'Right to be Forgotten'. However, many existing methods require access to the dat…

Towards Robust Evaluation of Unlearning in LLMs via Data Transformations

2024-11-23 · Abhinav Joshi, Shaswati Saha, Divyaksh Shukla, Sriram Vema 외

Large Language Models (LLMs) have shown to be a great success in a wide range of applications ranging from regular NLP-based use cases to AI agents. LLMs have been trained on a vast corpus of texts from various sources; …

Machine Unlearning