paper-with-me

Papers

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

2026-06-25 · William Poulett, Alice Waterhouse, Ben Wallace, Scarlett Kynoch, Amaia Imaz Blanco, Michael Spence, Jonathan Pearson arxiv

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

📄 PDF Abstract BibTeX arXiv:2606.26879

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Embedding-Driven Diversity Sampling to Improve Few-Shot Synthetic Data Generation

2025-01-20 · Ivan Lopez, Fateme Nateghi Haredasht, Kaitlin Caoili, Jonathan H Chen 외

Accurate classification of clinical text often requires fine-tuning pre-trained language models, a process that is costly and time-consuming due to the need for high-quality data and expert annotators. Synthetic data gen…

DiversitySynthetic Data Generation

ClinicalMamba: A Generative Clinical Language Model on Longitudinal Clinical Notes

2024-03-09 · Zhichao Yang, Avijit Mitra, Sunjae Kwon, Hong Yu

The advancement of natural language processing (NLP) systems in healthcare hinges on language model ability to interpret the intricate information contained within clinical notes. This process often requires integrating …

Few-Shot LearningLanguage ModelingLanguage ModellingMamba

DENSE: Longitudinal Progress Note Generation with Temporal Modeling of Heterogeneous Clinical Notes Across Hospital Visits

2025-07-18 · Garapati Keerthana, Manik Gupta

Progress notes are among the most clinically meaningful artifacts in an Electronic Health Record (EHR), offering temporally grounded insights into a patient's evolving condition, treatments, and care decisions. Despite t…

Large Language Model

NoteChat: A Dataset of Synthetic Doctor-Patient Conversations Conditioned on Clinical Notes

2023-10-24 · Junda Wang, Zonghai Yao, Zhichao Yang, Huixue Zhou 외

We introduce NoteChat, a novel cooperative multi-agent framework leveraging Large Language Models (LLMs) to generate patient-physician dialogues. NoteChat embodies the principle that an ensemble of role-specific LLMs, th…

Dialogue Generation

Investigating Alternative Feature Extraction Pipelines For Clinical Note Phenotyping

2023-10-05 · Neil Daniel

A common practice in the medical industry is the use of clinical notes, which consist of detailed patient observations. However, electronic health record systems frequently do not contain these observations in a structur…

Clinical Note Phenotyping