paper-with-me

홈 › Papers

Natural Identifiers for Privacy and Data Audits in Large Language Models

2026-06-23 · Lorenzo Rossi, Bartłomiej Marek, Franziska Boenisch, Adam Dziedzic arxiv

Assessing the privacy of large language models (LLMs) presents significant challenges. In particular, most existing methods for auditing differential privacy require the insertion of specially crafted canary data during training, making them impractical for auditing already-trained models without costly retraining. Additionally, dataset inference, which audits whether a suspect dataset was used to train a model, is infeasible without access to a private non-member held-out dataset. Yet, such held-out datasets are often unavailable or difficult to construct for real-world cases since they have to be from the same distribution (IID) as the suspect data. These limitations severely hinder the ability to conduct scalable, post-hoc audits. To enable such audits, this work introduces natural identifiers (NIDs) as a novel solution to the above-mentioned challenges. NIDs are structured random strings, such as cryptographic hashes and shortened URLs, naturally occurring in common LLM training datasets. Their format enables the generation of unlimited additional random strings from the same distribution, which can act as alternative canaries for audits and as same-distribution held-out data for dataset inference. Our evaluation highlights that indeed, using NIDs, we can facilitate post-hoc differential privacy auditing without any retraining and enable dataset inference for any suspect dataset containing NIDs without the need for a private non-member held-out dataset.

📄 PDF Abstract BibTeX arXiv:2606.24408

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

C3PA: An Open Dataset of Expert-Annotated and Regulation-Aware Privacy Policies to Enable Scalable Regulatory Compliance Audits

2024-10-04 · Maaz Bin Musa, Steven M. Winston, Garrison Allen, Jacob Schiller 외

The development of tools and techniques to analyze and extract organizations data habits from privacy policies are critical for scalable regulatory compliance audits. Unfortunately, these tools are becoming increasingly …

Membership Inference Attacks on Sequence Models

2025-06-05 · Lorenzo Rossi, Michael Aerni, Jie Zhang, Florian Tramèr

Sequence models, such as Large Language Models (LLMs) and autoregressive image generators, have a tendency to memorize and inadvertently leak sensitive information. While this tendency has critical legal implications, ex…

Inference AttackMembership Inference AttackMemorization

Human-Centred LLM Privacy Audits: Findings and Frictions

2026-03-12 · Dimitri Staufer, Kirsten Morehouse, David Hartmann, Bettina Berendt arxiv

Large language models (LLMs) learn statistical associations from massive training corpora and user interactions, and deployed systems can surface or infer information about individuals. Yet people lack practical ways to …

Disclosure Audits for LLM Agents

2025-06-11 · Saswat Das, Jameson Sandler, Ferdinando Fioretto

Large Language Model agents have begun to appear as personal assistants, customer service bots, and clinical aides. While these applications deliver substantial operational benefits, they also require continuous access t…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model

AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems

2026-02-12 · Faouzi El Yagoubi, Godwin Badu-Marfo, Ranwa Al Mallah arxiv

Multi-agent Large Language Model (LLM) systems create privacy risks that current output-only benchmarks cannot measure. When agents coordinate on tasks, sensitive data may pass through inter-agent messages, shared memory…