paper-with-me

Papers

Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks

2024-06-25 · Arij Riabi, Menel Mahamdi, Virginie Mouilleron, Djamé Seddah

Protecting privacy is essential when sharing data, particularly in the case of an online radicalization dataset that may contain personal information. In this paper, we explore the balance between preserving data usefulness and ensuring robust privacy safeguards, since regulations like the European GDPR shape how personal information must be handled. We share our method for manually pseudonymizing a multilingual radicalization dataset, ensuring performance comparable to the original data. Furthermore, we highlight the importance of establishing comprehensive guidelines for processing sensitive NLP data by sharing our complete pseudonymization process, our guidelines, the challenges we encountered as well as the resulting dataset.

📄 PDF Abstract BibTeX arXiv:2406.17875

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Similar Papers 제목 키워드 기반

Grandma Karl is 27 years old -- research agenda for pseudonymization of research data

2023-08-30 · Elena Volodina, Simon Dobnik, Therese Lindström Tiedemann, Xuan-Son Vu

Accessibility of research data is critical for advances in many research fields, but textual data often cannot be shared due to the personal and sensitive information which it contains, e.g names or political opinions. G…

Re-pseudonymization Strategies for Smart Meter Data Are Not Robust to Deep Learning Profiling Attacks

2024-04-05 · Ana-Maria Cretu, Miruna Rusu, Yves-Alexandre de Montjoye

Smart meters, devices measuring the electricity and gas consumption of a household, are currently being deployed at a fast rate throughout the world. The data they collect are extremely useful, including in the fight aga…

Say Something Else: Rethinking Contextual Privacy as Information Sufficiency

2026-04-07 · Yunze Xiao, Wenkai Li, Xiaoyuan Wu, Ningshan Ma 외 arxiv

LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private. Existing systems support only suppression (omitting sensitive informa…

Privacy- and Utility-Preserving NLP with Anonymized Data: A case study of Pseudonymization

2023-06-08 · Oleksandr Yermilov, Vipul Raheja, Artem Chernodub

This work investigates the effectiveness of different pseudonymization techniques, ranging from rule-based substitutions to using pre-trained Large Language Models (LLMs), on a variety of datasets and models used for two…

text-classificationText Classification

A General Pseudonymization Framework for Cloud-Based LLMs: Replacing Privacy Information in Controlled Text Generation

2025-02-21 · Shilong Hou, Ruilin Shang, Zi Long, Xianghua Fu 외

An increasing number of companies have begun providing services that leverage cloud-based large language models (LLMs), such as ChatGPT. However, this development raises substantial privacy concerns, as users' prompts ar…

Text Generation