paper-with-me

홈 › Papers

GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models

2025-09-25 · Jieli Zhu, Vi Ngoc-Nha Tran arxiv

Small language models (SLMs) become unprecedentedly appealing due to their approximately equivalent performance compared to large language models (LLMs) in certain fields with less energy and time consumption during training and inference. However, the personally identifiable information (PII) leakage of SLMs for downstream tasks has yet to be explored. In this study, we investigate the PII leakage of the chatbot based on SLM. We first finetune a new chatbot, i.e., ChatBioGPT based on the backbone of BioGPT using medical datasets Alpaca and HealthCareMagic. It shows a matchable performance in BERTscore compared with previous studies of ChatDoctor and ChatGPT. Based on this model, we prove that the previous template-based PII attacking methods cannot effectively extract the PII in the dataset for leakage detection under the SLM condition. We then propose GEP, which is a greedy coordinate gradient-based (GCG) method specifically designed for PII extraction. We conduct experimental studies of GEP and the results show an increment of up to 60$\times$ more leakage compared with the previous template-based methods. We further expand the capability of GEP in the case of a more complicated and realistic situation by conducting free-style insertion where the inserted PII in the dataset is in the form of various syntactic expressions instead of fixed templates, and GEP is still able to reveal a PII leakage rate of up to 4.53%.

📄 PDF Abstract BibTeX arXiv:2509.21192

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild

2024-07-16 · Niloofar Mireshghallah, Maria Antoniak, Yash More, Yejin Choi 외

Measuring personal disclosures made in human-chatbot interactions can provide a better understanding of users' AI literacy and facilitate privacy research for large language models (LLMs). We run an extensive, fine-grain…

Chatbot

Are Chatbots Ready for Privacy-Sensitive Applications? An Investigation into Input Regurgitation and Prompt-Induced Sanitization

2023-05-24 · Aman Priyanshu, Supriti Vijay, Ayush Kumar, Rakshit Naidu 외

LLM-powered chatbots are becoming widely adopted in applications such as healthcare, personal assistants, industry hiring decisions, etc. In many of these cases, chatbots are fed sensitive, personal information in their …

In-Context Learning

Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?

2023-07-31 · Albert Yu Sun, Eliott Zemour, Arushi Saxena, Udith Vaidyanathan 외

Machine learning practitioners often fine-tune generative pre-trained models like GPT-3 to improve model performance at specific tasks. Previous works, however, suggest that fine-tuned machine learning models memorize an…

Memorization

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models

2025-07-06 · Sathesh P. Sivashanmugam

Large language models (LLMs) have transformed natural language processing, but their ability to memorize training data poses significant privacy risks. This paper investigates model inversion attacks on the Llama 3.2 mod…

Privacy Preserving

Data Center Audio/Video Intelligence on Device (DAVID) -- An Edge-AI Platform for Smart-Toys

2023-11-18 · Gabriel Cosache, Francisco Salgado, Cosmin Rotariu, George Sterpu 외

An overview is given of the DAVID Smart-Toy platform, one of the first Edge AI platform designs to incorporate advanced low-power data processing by neural inference models co-located with the relevant image or audio sen…

text-to-speechText to Speech