A French Medical Conversations Corpus Annotated for a Virtual Patient Dialogue System
Data-driven approaches for creating virtual patient dialogue systems require the availability of large data specific to the language,domain and clinical cases studied. Based on the lack of dialogue corpora in French for medical education, we propose an annotatedcorpus of dialogues including medical consultation interactions between doctor and patient. In this work, we detail the building processof the proposed dialogue corpus, describe the annotation guidelines and also present the statistics of its contents. We then conducted aquestion categorization task to evaluate the benefits of the proposed corpus that is made publicly available.
Code (1)
Similar Papers 제목 키워드 기반
Ubuntu-fr: A Large and Open Corpus for Multi-modal Analysis of Online Written Conversations
We present a large, free, French corpus of online written conversations extracted from the Ubuntu platform{'}s forums, mailing lists and IRC channels. The corpus is meant to support multi-modality and diachronic studies …
FRASIMED: a Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection
Natural language processing (NLP) applications such as named entity recognition (NER) for low-resource corpora do not benefit from recent advances in the development of large language models (LLMs) where there is still a…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERThe Quaero French Medical Corpus: A Ressource for Medical Entity Recognition and Normalization
A vast amount of information in the biomedical domain is available as natural language free text. An increasing number of documents in the field are written in languages other than English. Therefore, it is essential to …
CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives
Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluating models. Such resources are scarce, especially for specialized domains in languages other than English. In par…
Semantic SimilaritySemantic Textual SimilaritySentenceSentence Embeddings+3AlloSat: A New Call Center French Corpus for Satisfaction and Frustration Analysis
We present a new corpus, named AlloSat, composed of real-life call center conversations in French that is continuously annotated in frustration and satisfaction. This corpus has been set up to develop new systems able to…