Role-Preserving Redaction of Medical Records to Enable Ontology-Driven Processing
Electronic medical records (EMR) have largely replaced hand-written patient files in healthcare. The growing pool of EMR data presents a significant resource in medical research, but the U.S. Health Insurance Portability and Accountability Act (HIPAA) mandates redacting medical records before performing any analysis on the same. This process complicates obtaining medical data and can remove much useful information from the record. As part of a larger project involving ontology-driven medical processing, we employ a method of recognizing protected health information (PHI) that maps to ontological terms. We then use the relationships defined in the ontology to redact medical texts so that roles and semantics of terms are retained without compromising anonymity. The method is evaluated by clinical experts on several hundred medical documents, achieving up to a 98.8{\%} f-score, and has already shown promise for retaining semantic information in later processing.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
RedactionBench
Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII). While redacting PII is a data cleaning prerequisite, existing benchmarks conflate e…
RedacBench: Can AI Erase Your Secrets?
Modern language models can readily extract sensitive information from unstructured text, making redaction -- the selective removal of such information -- critical for data security. However, existing benchmarks for redac…
Validating transformers for redaction of text from electronic health records in real-world healthcare
Protecting patient privacy in healthcare records is a top priority, and redaction is a commonly used method for obscuring directly identifiable information in text. Rule-based methods have been widely used, but their pre…
Selective Token-Level Cryptographic Redaction for Privacy-Preserving Clinical Deployment of Large Language Models
While large language models (LLMs) are increasingly used for clinical applications, many existing pipelines require sending raw sensitive health information to remote servers for processing, which heightens the risk of p…
Question AnsweringFrom Redaction to Restoration: Deep Learning for Medical Image Anonymization and Reconstruction
Removing patient-specific information from medical images is crucial to enable sharing and open science without compromising patient identities. However, many methods currently used for deidentification have negative eff…
Image Reconstruction