HB Deid - HB De-identification tool demonstrator
This paper describes a freely available web-based demonstrator called HB Deid. HB Deid identifies so-called protected health information, PHI, in a text written in Swedish and removes, masks, or replaces them with surrogates or pseudonyms. PHIs are named entities such as personal names, locations, ages, phone numbers, dates. HB Deid uses a CRF model trained on non-sensitive annotated text in Swedish, as well as a rule-based post-processing step for finding PHI. The final step in obscuring the PHI is then to either mask it, show only the class name or use a rule-based pseudonymisation system to replace it.
Code (0)
등록된 구현이 없습니다.
Tasks
De-identificationSimilar Papers 제목 키워드 기반
Pyclipse, a library for deidentification of free-text clinical notes
Automated deidentification of clinical text data is crucial due to the high cost of manual deidentification, which has been a barrier to sharing clinical text and the advancement of clinical natural language processing. …
Face Deidentification with Generative Deep Neural Networks
Face deidentification is an active topic amongst privacy and security researchers. Early deidentification methods relying on image blurring or pixelization were replaced in recent years with techniques based on formal an…
FDeID-Toolbox: Face De-Identification Toolbox
Face de-identification (FDeID) aims to remove personally identifiable information from facial images while preserving task-relevant utility attributes such as age, gender, and expression. It is critical for privacy-prese…
Age EstimationDIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization
Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporar…
De-identificationLanguage ModelingLanguage ModellingLarge Language Model+1Diverse Community Data for Benchmarking Data Privacy Algorithms
The Collaborative Research Cycle (CRC) is a National Institute of Standards and Technology (NIST) benchmarking program intended to strengthen understanding of tabular data deidentification technologies. Deidentification …
Benchmarking