paper-with-me

홈 › Papers

Reconstruction of Personally Identifiable Information from Supervised Finetuned Models

2026-05-12 · Sae Furukawa, Alina Oprea arxiv

Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to domain-specific, instruction-following tasks. SFT datasets, composed of instruction-response pairs, often include user-provided information that may contain sensitive data such as personally identifiable information (PII), raising privacy concerns. This paper studies the problem of PII reconstruction from SFT models for the first time. We construct multi-turn, user-centric Q&A datasets in sensitive domains, specifically medical and legal settings, that incorporate PII to enable realistic evaluation of leakage. Using these datasets, we evaluate the extent to which an adversary, with varying levels of knowledge about the fine-tuning dataset, can infer sensitive information about individuals whose data was used during SFT. In the reconstruction setting, we propose COVA, a novel decoding algorithm to reconstruct PII under prefix-based attacks, consistently outperforming existing extraction methods. Our results show that even partial attacker knowledge can significantly improve reconstruction success, while leakage varies substantially across PII types.

📄 PDF Abstract BibTeX arXiv:2605.12264

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Conundrum of Trustworthy Research on Attacking Personally Identifiable Information Removal Techniques

2026-03-09 · Sebastian Ochs, Ivan Habernal arxiv

Removing personally identifiable information (PII) from texts is necessary to comply with various data protection regulations and to enable data sharing without compromising privacy. However, recent works show that docum…

Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework

2026-01-07 · Xiaoyu Luo, Yiyi Chen, Qiongxiu Li, Johannes Bjerva arxiv

Large Language Models (LLMs) have been reported to "leak" Personally Identifiable Information (PII), with successful PII reconstruction often interpreted as evidence of memorization. We propose a principled revision of m…

Analyzing Leakage of Personally Identifiable Information in Language Models

2023-02-01 · Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople 외

Language Models (LMs) have been shown to leak information about training data through sentence-level membership inference and reconstruction attacks. Understanding the risk of LMs leaking Personally Identifiable Informat…

Sentence

Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?

2023-07-31 · Albert Yu Sun, Eliott Zemour, Arushi Saxena, Udith Vaidyanathan 외

Machine learning practitioners often fine-tune generative pre-trained models like GPT-3 to improve model performance at specific tasks. Previous works, however, suggest that fine-tuned machine learning models memorize an…

Memorization

GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition

2023-01-05 · NeurIPS 2023 11

Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personal…

ObjectObject Recognition