paper-with-me

Papers

HitzalMed: Anonymisation of Clinical Text in Spanish

2020-05-01 · LREC 2020 5 · Salvador Lima Lopez, Naiara Perez, Laura Garc{\'\i}a-Sardi{\~n}a, Montse Cuadros

HitzalMed is a web-framed tool that performs automatic detection of sensitive information in clinical texts using machine learning algorithms reported to be competitive for the task. Moreover, once sensitive information is detected, different anonymisation techniques are implemented that are configurable by the user {--}for instance, substitution, where sensitive items are replaced by same category text in an effort to generate a new document that looks as natural as the original one. The tool is able to get data from different document formats and outputs downloadable anonymised data. This paper presents the anonymisation and substitution technology and the demonstrator which is publicly available at https://snlt.vicomtech.org/hitzalmed.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT

2020-03-06 · LREC 2020 5 · Aitor García-Pablos, Naiara Perez, Montse Cuadros

Massive digital data processing provides a wide range of opportunities and benefits, but at the cost of endangering personal data privacy. Anonymisation consists in removing or replacing sensitive information from data, …

Feature EngineeringGeneral Classification

Spanish Datasets for Sensitive Entity Detection in the Legal Domain

2022-06-01 · LREC 2022 6 · Ona de Gibert Bonet, Aitor García Pablos, Montse Cuadros, Maite Melero

The de-identification of sensible data, also known as automatic textual anonymisation, is essential for data sharing and reuse, both for research and commercial purposes. The first step for data anonymisation is the dete…

De-identification

Anonymisation Models for Text Data: State of the art, Challenges and Future Directions

2021-08-01 · ACL 2021 5 · Pierre Lison, Ildik{\'o} Pil{\'a}n, David Sanchez, Montserrat Batet 외

This position paper investigates the problem of automated text anonymisation, which is a prerequisite for secure sharing of documents containing sensitive information about individuals. We summarise the key concepts behi…

PositionPrivacy Preserving

Textwash -- automated open-source text anonymisation

2022-08-27 · Bennett Kleinberg, Toby Davies, Maximilian Mozes

The increased use of text data in social science research has benefited from easy-to-access data (e.g., Twitter). That trend comes at the cost of research requiring sensitive but hard-to-share data (e.g., interview data,…

Annotation of negation in the IULA Spanish Clinical Record Corpus

2017-04-01 · WS 2017 4 · Montserrat Marimon, Jorge Vivaldi, N{\'u}ria Bel

This paper presents the IULA Spanish Clinical Record Corpus, a corpus of 3,194 sentences extracted from anonymized clinical records and manually annotated with negation markers and their scope. The corpus was conceived a…

Medical DiagnosisNegationNegation DetectionTerm Extraction