paper-with-me

Papers

Textwash -- automated open-source text anonymisation

2022-08-27 · Bennett Kleinberg, Toby Davies, Maximilian Mozes

The increased use of text data in social science research has benefited from easy-to-access data (e.g., Twitter). That trend comes at the cost of research requiring sensitive but hard-to-share data (e.g., interview data, police reports, electronic health records). We introduce a solution to that stalemate with the open-source text anonymisation software_Textwash_. This paper presents the empirical evaluation of the tool using the TILD criteria: a technical evaluation (how accurate is the tool?), an information loss evaluation (how much information is lost in the anonymisation process?) and a de-anonymisation test (can humans identify individuals from anonymised text data?). The findings suggest that Textwash performs similar to state-of-the-art entity recognition models and introduces a negligible information loss of 0.84%. For the de-anonymisation test, we tasked humans to identify individuals by name from a dataset of crowdsourced person descriptions of very famous, semi-famous and non-existing individuals. The de-anonymisation rate ranged from 1.01-2.01% for the realistic use cases of the tool. We replicated the findings in a second study and concluded that Textwash succeeds in removing potentially sensitive information that renders detailed person descriptions practically anonymous.

📄 PDF Abstract BibTeX arXiv:2208.13081

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Anonymisation Models for Text Data: State of the art, Challenges and Future Directions

2021-08-01 · ACL 2021 5 · Pierre Lison, Ildik{\'o} Pil{\'a}n, David Sanchez, Montserrat Batet 외

This position paper investigates the problem of automated text anonymisation, which is a prerequisite for secure sharing of documents containing sensitive information about individuals. We summarise the key concepts behi…

PositionPrivacy Preserving

The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation

2024-07-16 · Michele Panariello, Natalia Tomashenko, Xin Wang, Xiaoxiao Miao 외

The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the …

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionVoice Conversion

AnonySIGN: Novel Human Appearance Synthesis for Sign Language Video Anonymisation

2021-07-22 · Ben Saunders, Necati Cihan Camgoz, Richard Bowden

The visual anonymisation of sign language data is an essential task to address privacy concerns raised by large-scale dataset collection. Previous anonymisation techniques have either significantly affected sign comprehe…

Image-to-Image Translation

The Multilingual Anonymisation Toolkit for Public Administrations (MAPA) Project

2020-11-01 · EAMT 2020 11 · Ēriks Ajausks, Victoria Arranz, Laurent Bié, Aleix Cerdà-i-Cucó 외

We describe the MAPA project, funded under the Connecting Europe Facility programme, whose goal is the development of an open-source de-identification toolkit for all official European Union languages. It will be develop…

De-identification

Benchmarking Advanced Text Anonymisation Methods: A Comparative Study on Novel and Traditional Approaches

2024-04-22 · Dimitris Asimopoulos, Ilias Siniosoglou, Vasileios Argyriou, Thomai Karamitsou 외

In the realm of data privacy, the ability to effectively anonymise text is paramount. With the proliferation of deep learning and, in particular, transformer architectures, there is a burgeoning interest in leveraging th…

BenchmarkingDiversity