paper-with-me

홈 › Papers

Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs

2024-07-08 · Sanjeet Singh, Shreya Gupta, Niralee Gupta, Naimish Sharma, Lokesh Srivastava, Vibhu Agarwal, Ashutosh Modi

The consequences of a healthcare data breach can be devastating for the patients, providers, and payers. The average financial impact of a data breach in recent months has been estimated to be close to USD 10 million. This is especially significant for healthcare organizations in India that are managing rapid digitization while still establishing data governance procedures that align with the letter and spirit of the law. Computer-based systems for de-identification of personal information are vulnerable to data drift, often rendering them ineffective in cross-institution settings. Therefore, a rigorous assessment of existing de-identification against local health datasets is imperative to support the safe adoption of digital health initiatives in India. Using a small set of de-identified patient discharge summaries provided by an Indian healthcare institution, in this paper, we report the nominal performance of de-identification algorithms (based on language models) trained on publicly available non-Indian datasets, pointing towards a lack of cross-institutional generalization. Similarly, experimentation with off-the-shelf de-identification systems reveals potential risks associated with the approach. To overcome data scarcity, we explore generating synthetic clinical reports (using publicly available and Indian summaries) by performing in-context learning over Large Language Models (LLMs). Our experiments demonstrate the use of generated reports as an effective strategy for creating high-performing de-identification systems with good generalization capabilities.

📄 PDF Abstract BibTeX arXiv:2407.05887

Code (1)

exploration-lab/llm-for-clinical-report-generation-deidentification 공식 구현

Tasks

De-identificationIn-Context Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

De-identifying Australian Hospital Discharge Summaries: An End-to-End Framework using Ensemble of Deep Learning Models

2021-01-01 · Leibo Liu, Oscar Perez-Concha, Anthony Nguyen, Vicki Bennett 외

Electronic Medical Records (EMRs) contain clinical narrative text that is of great potential value to medical researchers. However, this information is mixed with Personally Identifiable Information (PII) that presents r…

De-identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

2026-06-14 · Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang 외 arxiv

Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision…

Question Answering

LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review

2025-07-07 · Cheng Yuan, Xinkai Rui, Yongqi Fan, Yawei Fan 외 arxiv

Despite the remarkable performance of Large Language Models (LLMs) in automated discharge summary generation, they still suffer from hallucination issues, such as generating inaccurate content or fabricating information …

Overview of the First Shared Task on Clinical Text Generation: RRG24 and "Discharge Me!"

2024-09-25 · Justin Xu, Zhihong Chen, Andrew Johnston, Louis Blankemeier 외

Recent developments in natural language generation have tremendous implications for healthcare. For instance, state-of-the-art systems could automate the generation of sections in clinical reports to alleviate physician …

Text Generation

Rare Disease Identification from Clinical Notes with Ontologies and Weak Supervision

2021-05-05 · Hang Dong, Víctor Suárez-Paniagua, Huayu Zhang, Minhong Wang 외

The identification of rare diseases from clinical notes with Natural Language Processing (NLP) is challenging due to the few cases available for machine learning and the need of data annotation from clinical experts. We …

Entity Linking