paper-with-me

Papers

Unsupervised Text Deidentification

2022-10-20 · John X. Morris, Justin T. Chiu, Ramin Zabih, Alexander M. Rush

Deidentification seeks to anonymize textual data prior to distribution. Automatic deidentification primarily uses supervised named entity recognition from human-labeled data points. We propose an unsupervised deidentification method that masks words that leak personally-identifying information. The approach utilizes a specially trained reidentification model to identify individuals from redacted personal documents. Motivated by K-anonymity based privacy, we generate redactions that ensure a minimum reidentification rank for the correct profile of the document. To evaluate this approach, we consider the task of deidentifying Wikipedia Biographies, and evaluate using an adversarial reidentification metric. Compared to a set of unsupervised baselines, our approach deidentifies documents more completely while removing fewer words. Qualitatively, we see that the approach eliminates many identifying aspects that would fall outside of the common named entity based approach.

📄 PDF Abstract BibTeX arXiv:2210.11528

Code (1)

jxmorris12/unsupervised-text-deidentification 공식 구현 pytorch

Tasks

Named Entity RecognitionNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

Pyclipse, a library for deidentification of free-text clinical notes

2023-11-05 · Callandra Moore, Jonathan Ranisau, Walter Nelson, Jeremy Petch 외

Automated deidentification of clinical text data is crucial due to the high cost of manual deidentification, which has been a barrier to sharing clinical text and the advancement of clinical natural language processing. …

Face Deidentification with Generative Deep Neural Networks

2017-07-28 · Blaž Meden, Refik Can Malli, Sebastjan Fabijan, Hazim Kemal Ekenel 외

Face deidentification is an active topic amongst privacy and security researchers. Early deidentification methods relying on image blurring or pixelization were replaced in recent years with techniques based on formal an…

DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization

2024-10-22 · John X. Morris, Thomas R. Campion, Sri Laasya Nutheti, Yifan Peng 외

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporar…

De-identificationLanguage ModelingLanguage ModellingLarge Language Model+1

Diverse Community Data for Benchmarking Data Privacy Algorithms

2023-06-20 · NeurIPS 2023 11

The Collaborative Research Cycle (CRC) is a National Institute of Standards and Technology (NIST) benchmarking program intended to strengthen understanding of tabular data deidentification technologies. Deidentification …

Benchmarking

Speaker Identification Experiments Under Gender De-Identification

2022-03-09 · Marcos Faundez-Zanuy, Enric Sesa-Nogueras, Stefano Marinozzi

The present work is based on the COST Action IC1206 for De-identification in multimedia content. It was performed to test four algorithms of voice modifications on a speech gender recognizer to find the degree of modific…

De-identificationSpeaker Identification