paper-with-me

홈 › Papers

FRASIMED: a Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection

2023-09-19 · Jamil Zaghir, Mina Bjelogrlic, Jean-Philippe Goldman, Soukaïna Aananou, Christophe Gaudet-Blavignac, Christian Lovis

Natural language processing (NLP) applications such as named entity recognition (NER) for low-resource corpora do not benefit from recent advances in the development of large language models (LLMs) where there is still a need for larger annotated datasets. This research article introduces a methodology for generating translated versions of annotated datasets through crosslingual annotation projection. Leveraging a language agnostic BERT-based approach, it is an efficient solution to increase low-resource corpora with few human efforts and by only using already available open data resources. Quantitative and qualitative evaluations are often lacking when it comes to evaluating the quality and effectiveness of semi-automatic data generation strategies. The evaluation of our crosslingual annotation projection approach showed both effectiveness and high accuracy in the resulting dataset. As a practical application of this methodology, we present the creation of French Annotated Resource with Semantic Information for Medical Entities Detection (FRASIMED), an annotated corpus comprising 2'051 synthetic clinical cases in French. The corpus is now available for researchers and practitioners to develop and refine French natural language processing (NLP) applications in the clinical field (https://zenodo.org/record/8355629), making it the largest open annotated corpus with linked medical concepts in French.

📄 PDF Abstract BibTeX arXiv:2309.10770

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

FRACCO: A gold-standard annotated corpus of oncological entities with ICD-O-3.1 normalisation

2025-10-13 · Johann Pignat, Milena Vucetic, Christophe Gaudet-Blavignac, Jamil Zaghir 외 arxiv

Developing natural language processing tools for clinical text requires annotated datasets, yet French oncology resources remain scarce. We present FRACCO (FRench Annotated Corpus for Clinical Oncology) an expert-annotat…

CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives

2022-06-01 · LREC 2022 6 · Nicolas Hiebel, Olivier Ferret, Karën Fort, Aurélie Névéol

Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluating models. Such resources are scarce, especially for specialized domains in languages other than English. In par…

Semantic SimilaritySemantic Textual SimilaritySentenceSentence Embeddings+3

Annotation of specialized corpora using a comprehensive entity and relation scheme

2014-05-01 · LREC 2014 5 · Louise Del{\'e}ger, Anne-Laure Ligozat, Cyril Grouin, Pierre Zweigenbaum 외

Annotated corpora are essential resources for many applications in Natural Language Processing. They provide insight on the linguistic and semantic characteristics of the genre and domain covered, and can be used for the…

Relation

E:Calm Resource: a Resource for Studying Texts Produced by French Pupils and Students

2020-05-01 · LREC 2020 5 · Lydia-Mai Ho-Dac, Serge Fleury, Claude Ponton

The E:Calm resource is constructed from French student texts produced in a variety of usual contexts of teaching. The distinction of the E:Calm resource is to provide an ecological data set that gives a broad overview of…

POSPOS Tagging

Impact of translation on biomedical information extraction from real-life clinical notes

2023-06-03 · Christel Gérardin, Yuhan Xiong, Perceval Wajsbürt, Fabrice Carrat 외

The objective of our study is to determine whether using English tools to extract and normalize French medical concepts on translations provides comparable performance to French models trained on a set of annotated Frenc…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1