Anonymisation Models for Text Data: State of the art, Challenges and Future Directions
This position paper investigates the problem of automated text anonymisation, which is a prerequisite for secure sharing of documents containing sensitive information about individuals. We summarise the key concepts behind text anonymisation and provide a review of current approaches. Anonymisation methods have so far been developed in two fields with little mutual interaction, namely natural language processing and privacy-preserving data publishing. Based on a case study, we outline the benefits and limitations of these approaches and discuss a number of open challenges, such as (1) how to account for multiple types of semantic inferences, (2) how to strike a balance between disclosure risk and data utility and (3) how to evaluate the quality of the resulting anonymisation. We lay out a case for moving beyond sequence labelling models and incorporate explicit measures of disclosure risk into the text anonymisation process.
Code (1)
Tasks
PositionPrivacy PreservingSimilar Papers 제목 키워드 기반
Benchmarking Advanced Text Anonymisation Methods: A Comparative Study on Novel and Traditional Approaches
In the realm of data privacy, the ability to effectively anonymise text is paramount. With the proliferation of deep learning and, in particular, transformer architectures, there is a burgeoning interest in leveraging th…
BenchmarkingDiversityVocoder drift compensation by x-vector alignment in speaker anonymisation
For the most popular x-vector-based approaches to speaker anonymisation, the bulk of the anonymisation can stem from vocoding rather than from the core anonymisation function which is used to substitute an original speak…
Textwash -- automated open-source text anonymisation
The increased use of text data in social science research has benefited from easy-to-access data (e.g., Twitter). That trend comes at the cost of research requiring sensitive but hard-to-share data (e.g., interview data,…
The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the …
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionVoice ConversionEvaluating the Efficacy of AI Techniques in Textual Anonymization: A Comparative Study
In the digital era, with escalating privacy concerns, it's imperative to devise robust strategies that protect private data while maintaining the intrinsic value of textual information. This research embarks on a compreh…