paper-with-me

Papers

eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

2026-06-10 · Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini arxiv

We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. The corpus, in its current version, is composed of approximately 4 million clinical notes fully anonymized, covering diverse phases of patient care during the stay in the emergency department. In addition, a subset of about six thousand notes has been manually annotated by clinical experts through a structured Case Report Form (CRF) containing 132 items relevant for two patient situations in emergency departments, dyspnea and loss of consciousness. Items may assume numerical values (e.g., for blood saturation), categorical (e.g., for level of consciousness ), binary (e.g., for presence of traumas), and mixed value types. The annotation process involved multiple clinicians and underwent iterative revision to resolve ambiguities in item formulation, resulting in a richly structured (although high imbalanced) resource. The dataset aims to fill a relevant gap of data able to support both the development and the use of Large Language Models in concrete medical applications. We describe the data collection protocol, the on-site anonymisation pipeline, corpus statistics, and the annotation scheme. Finally, we propose CRF-filling as a novel structured information extraction benchmark, and provide zero-shot baseline resulting from Gemma-27B and MedGemma-27B. To the best of our knowledge, eCREAM-MedCorpus is the largest freely available dataset of clinical notes existing for the Italian language.

📄 PDF Abstract BibTeX arXiv:2606.12569

Code (0)

등록된 구현이 없습니다.

Tasks

Information Extraction

Similar Papers 제목 키워드 기반

Beyond Single-Feature Importance with ICECREAM

2023-07-19 · Michael Oesterle, Patrick Blöbaum, Atalanti A. Mastakouri, Elke Kirschbaum

Which set of features was responsible for a certain output of a machine learning model? Which components caused the failure of a cloud computing application? These are just two examples of questions we are addressing in …

Cloud ComputingFeature Importance

Demo: Guide-RAG: Evidence-Driven Corpus Curation for Retrieval-Augmented Generation in Long COVID

2025-10-17 · Philip DiGiacomo, Haoyang Wang, Jinrui Fang, Yan Leng 외 arxiv

As AI chatbots gain adoption in clinical medicine, developing effective frameworks for complex, emerging diseases presents significant challenges. We developed and evaluated six Retrieval-Augmented Generation (RAG) corpu…

Question Answering

The Leaf Clinical Trials Corpus: a new resource for query generation from clinical trial eligibility criteria

2022-07-27 · Nicholas J Dobbins, Tony Mullen, Ozlem Uzuner, Meliha Yetisgen

Identifying cohorts of patients based on eligibility criteria such as medical conditions, procedures, and medication use is critical to recruitment for clinical trials. Such criteria are often most naturally described in…

emrQA: A Large Corpus for Question Answering on Electronic Medical Records

2018-09-03 · EMNLP 2018 10 · Anusri Pampari, Preethi Raghavan, Jennifer Liang, Jian Peng

We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generati…

FormQuestion Answering

Annotation of a Large Clinical Entity Corpus

2018-10-01 · EMNLP 2018 10 · Pinal Patel, Disha Davey, Vishal Panchal, Parth Pathak

Having an entity annotated corpus of the clinical domain is one of the basic requirements for detection of clinical entities using machine learning (ML) approaches. Past researches have shown the superiority of statistic…

Machine TranslationSmall Data Image Classification