paper-with-me

홈 › Papers

FRACCO: A gold-standard annotated corpus of oncological entities with ICD-O-3.1 normalisation

2025-10-13 · Johann Pignat, Milena Vucetic, Christophe Gaudet-Blavignac, Jamil Zaghir, Amandine Stettler, Fanny Amrein, Jonatan Bonjour, Jean-Philippe Goldman, Olivier Michielin, Christian Lovis, Mina Bjelogrlic arxiv

Developing natural language processing tools for clinical text requires annotated datasets, yet French oncology resources remain scarce. We present FRACCO (FRench Annotated Corpus for Clinical Oncology) an expert-annotated corpus of 1301 synthetic French clinical cases, initially translated from the Spanish CANTEMIST corpus as part of the FRASIMED initiative. Each document is annotated with terms related to morphology, topography, and histologic differentiation, using the International Classification of Diseases for Oncology (ICD-O) as reference. An additional annotation layer captures composite expression-level normalisations that combine multiple ICD-O elements into unified clinical concepts. Annotation quality was ensured through expert review: 1301 texts were manually annotated for entity spans by two domain experts. A total of 71127 ICD-O normalisations were produced through a combination of automated matching and manual validation by a team of five annotators. The final dataset representing 399 unique morphology codes (from 2549 different expressions), 272 topography codes (from 3143 different expressions), and 2043 unique composite expressions (from 11144 different expressions). This dataset provides a reference standard for named entity recognition and concept normalisation in French oncology texts.

📄 PDF Abstract BibTeX arXiv:2510.13873

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning

2025-10-21 · Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe 외 arxiv

Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we…

Semantic Similarity

CALBC: Releasing the Final Corpora

2012-05-01 · LREC 2012 5 · {\c{S}}enay Kafkas, Ian Lewin, David Milward, Erik van Mulligen 외

A number of gold standard corpora for named entity recognition are available to the public. However, the existing gold standard corpora are limited in size and semantic entity types. These usually lead to implementation …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

MoNERo: a Biomedical Gold Standard Corpus for the Romanian Language

2019-08-01 · WS 2019 8 · Maria Mitrofan, Verginica Barbu Mititelu, Grigorina Mitrofan

In an era when large amounts of data are generated daily in various fields, the biomedical field among others, linguistic resources can be exploited for various tasks of Natural Language Processing. Moreover, increasing …

Phrase Detectives Corpus 1.0 Crowdsourced Anaphoric Coreference.

2016-05-01 · LREC 2016 5 · Jon Chamberlain, Massimo Poesio, Udo Kruschwitz

Natural Language Engineering tasks require large and complex annotated datasets to build more advanced models of language. Corpora are typically annotated by several experts to create a gold standard; however, there are …

text annotation

Turning silver into gold: error-focused corpus reannotation with active learning

2019-09-01 · RANLP 2019 9 · Pierre Andr{\'e} M{\'e}nard, Antoine Mougeot

While high quality gold standard annotated corpora are crucial for most tasks in natural language processing, many annotated corpora published in recent years, created by annotators or tools, contains noisy annotations. …

Active LearningDocument ClassificationPart-Of-Speech Tagging