paper-with-me

Papers

PARHAF, a human-authored corpus of clinical reports for fictitious patients in French

2026-03-20 · Xavier Tannier, Salam Abbara, Rémi Flicoteaux, Youness Khalil, Aurélie Névéol, Pierre Zweigenbaum, Emmanuel Bacry arxiv

The development of clinical natural language processing (NLP) systems is severely hampered by the sensitive nature of medical records, which restricts data sharing under stringent privacy regulations, particularly in France and the broader European Union. To address this gap, we introduce PARHAF, a large open-source corpus of clinical documents in French. PARHAF comprises expert-authored clinical reports describing realistic yet entirely fictitious patient cases, making it anonymous and freely shareable by design. The corpus was developed using a structured protocol that combined clinician expertise with epidemiological guidance from the French National Health Data System (SNDS), ensuring broad clinical coverage. A total of 104 medical residents across 18 specialties authored and peer-reviewed the reports following predefined clinical scenarios and document templates. The corpus contains 7394 clinical reports covering 5009 patient cases across a wide range of medical and surgical specialties. It includes a general-purpose component designed to approximate real-world hospitalization distributions, and four specialized subsets that support information-extraction use cases in oncology, infectious diseases, and diagnostic coding. Documents are released under a CC-BY open license, with a portion temporarily embargoed to enable future benchmarking under controlled conditions. PARHAF provides a valuable resource for training and evaluating French clinical language models in a fully privacy-preserving setting, and establishes a replicable methodology for building shareable synthetic clinical corpora in other languages and health systems.

📄 PDF Abstract BibTeX arXiv:2603.20494

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PARROT: An Open Multilingual Radiology Reports Dataset

2025-07-25 · Bastien Le Guellec, Kokou Adambounou, Lisa C Adams, Thibault Agripnidis 외 arxiv

Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages f…

Using Large Language Models To Translate Machine Results To Human Results

2025-12-30 · Trishna Niraula, Jonathan Stubblefield arxiv

Artificial intelligence (AI) has transformed medical imaging, with computer vision (CV) systems achieving state-of-the-art performance in classification and detection tasks. However, these systems typically output struct…

Semantic SimilarityAnomaly Detection

ER-REASON: A Benchmark Dataset for LLM-Based Clinical Reasoning in the Emergency Room

2025-05-28 · Nikita Mehandru, Niloufar Golchini, David Bamman, Travis Zack 외

Large language models (LLMs) have been extensively evaluated on medical question answering tasks based on licensing exams. However, real-world evaluations often depend on costly human annotators, and existing benchmarks …

Medical Question AnsweringQuestion Answering

Annotating Negation in Spanish Clinical Texts

2017-04-01 · WS 2017 4 · Noa Cruz, Roser Morante, Manuel J. Ma{\~n}a L{\'o}pez, Jacinto Mata V{\'a}zquez 외

In this paper we present on-going work on annotating negation in Spanish clinical documents. A corpus of anamnesis and radiology reports has been annotated by two domain expert annotators with negation markers and negate…

Negation

Introducing a Corpus of Human-Authored Dialogue Summaries in Portuguese

2013-09-01 · RANLP 2013 9 · Norton Trevisan Roman, Paul Piwek, Ariadne M. B. Rizzoni Carvalho, Alex Rossi Alvares 외