Development and validation of a natural language processing algorithm to pseudonymize documents in the context of a clinical data warehouse
The objective of this study is to address the critical issue of de-identification of clinical reports in order to allow access to data for research purposes, while ensuring patient privacy. The study highlights the difficulties faced in sharing tools and resources in this domain and presents the experience of the Greater Paris University Hospitals (AP-HP) in implementing a systematic pseudonymization of text documents from its Clinical Data Warehouse. We annotated a corpus of clinical documents according to 12 types of identifying entities, and built a hybrid system, merging the results of a deep learning model as well as manual rules. Our results show an overall performance of 0.99 of F1-score. We discuss implementation choices and present experiments to better understand the effort involved in such a task, including dataset size, document types, language models, or rule addition. We share guidelines and code under a 3-Clause BSD license.
Code (0)
등록된 구현이 없습니다.
Tasks
De-identificationSimilar Papers 제목 키워드 기반
Development and Validation of MicrobEx: an Open-Source Package for Microbiology Culture Concept Extraction
Microbiology culture reports contain critical information for important clinical and public health applications. However, microbiology reports often have complex, semi-structured, free-text data that present a barrier fo…
Cultural Vocal Bursts Intensity PredictionTowards an Automated Requirements-driven Development of Smart Cyber-Physical Systems
The Invariant Refinement Method for Self Adaptation (IRM-SA) is a design method targeting development of smart Cyber-Physical Systems (sCPS). It allows for a systematic translation of the system requirements into the sys…
TranslationA Natural Language Processing Framework for Hotel Recommendation Based on Users' Text Reviews
Recently, the application of Artificial Intelligence algorithms in hotel recommendation systems has become an increasingly popular topic. One such method that has proven to be effective in this field is Deep Learning, es…
Recommendation SystemsClusterDataSplit: Exploring Challenging Clustering-Based Data Splits for Model Performance Evaluation
This paper adds to the ongoing discussion in the natural language processing community on how to choose a good development set. Motivated by the real-life necessity of applying machine learning models to different data d…
ClusteringPatent classificationSentiment AnalysisAutomated Fact-Checking: A Survey
As online false information continues to grow, automated fact-checking has gained an increasing amount of attention in recent years. Researchers in the field of Natural Language Processing (NLP) have contributed to the t…
Fact CheckingSurvey