paper-with-me

Papers

PII-Scope: A Benchmark for Training Data PII Leakage Assessment in LLMs

2024-10-09 · Krishna Kanth Nakka, Ahmed Frikha, Ricardo Mendes, Xue Jiang, Xuebing Zhou

In this work, we introduce PII-Scope, a comprehensive benchmark designed to evaluate state-of-the-art methodologies for PII extraction attacks targeting LLMs across diverse threat settings. Our study provides a deeper understanding of these attacks by uncovering several hyperparameters (e.g., demonstration selection) crucial to their effectiveness. Building on this understanding, we extend our study to more realistic attack scenarios, exploring PII attacks that employ advanced adversarial strategies, including repeated and diverse querying, and leveraging iterative learning for continual PII extraction. Through extensive experimentation, our results reveal a notable underestimation of PII leakage in existing single-query attacks. In fact, we show that with sophisticated adversarial capabilities and a limited query budget, PII extraction rates can increase by up to fivefold when targeting the pretrained model. Moreover, we evaluate PII leakage on finetuned models, showing that they are more vulnerable to leakage than pretrained models. Overall, our work establishes a rigorous empirical benchmark for PII extraction attacks in realistic threat scenarios and provides a strong foundation for developing effective mitigation strategies.

📄 PDF Abstract BibTeX arXiv:2410.06704

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-P$^2$A: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models

2024-12-27 · Jie Zhang, Xiangkui Cao, Zhouyu Han, Shiguang Shan 외

Large Vision-Language Models (LVLMs) exhibit impressive potential across various tasks but also face significant privacy risks, limiting their practical applications. Current researches on privacy assessment for LVLMs is…

(Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs

2025-10-07 · Jiashu Tao, Reza Shokri arxiv

Machine learning models are known to leak sensitive information, as they inevitably memorize (parts of) their training data. More alarmingly, large language models (LLMs) are now trained on nearly all available data, whi…

Computational Efficiency

A Grammar of Machine Learning Workflows: Rejecting Data Leakage at Call Time

2026-03-11 · Simon Roth arxiv

Data leakage has been identified in 648 published papers across 30 scientific fields. The knowledge to prevent it has existed for over a decade; the problem persists because the tools do not enforce what the textbooks te…

Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds

2026-08-27 · Waqas Khan, Tabinda Sarwar, Jingyue Cong, Xun Yi 외 arxiv

Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligations. Selective unlearning removes requested knowledge while preserving model uti…

AI Idea Bench 2025: AI Research Idea Generation Benchmark

2025-04-19 · Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li 외

Large-scale Language Models (LLMs) have revolutionized human-AI interaction and achieved significant success in the generation of novel ideas. However, current assessments of idea generation overlook crucial factors such…

Benchmarkingscientific discovery