paper-with-me

홈 › Papers

Multi-label Annotation in Scientific Articles - The Multi-label Cancer Risk Assessment Corpus

2016-05-01 · LREC 2016 5 · James Ravenscroft, Anika Oellrich, Shyamasree Saha, Maria Liakata

With the constant growth of the scientific literature, automated processes to enable access to its contents are increasingly in demand. Several functional discourse annotation schemes have been proposed to facilitate information extraction and summarisation from scientific articles, the most well known being argumentative zoning. Core Scientific concepts (CoreSC) is a three layered fine-grained annotation scheme providing content-based annotations at the sentence level and has been used to index, extract and summarise scientific publications in the biomedical literature. A previously developed CoreSC corpus on which existing automated tools have been trained contains a single annotation for each sentence. However, it is the case that more than one CoreSC concept can appear in the same sentence. Here, we present the Multi-CoreSC CRA corpus, a text corpus specific to the domain of cancer risk assessment (CRA), consisting of 50 full text papers, each of which contains sentences annotated with one or more CoreSCs. The full text papers have been annotated by three biology experts. We present several inter-annotator agreement measures appropriate for multi-label annotation assessment. Employing several inter-annotator agreement measures, we were able to identify the most reliable annotator and we built a harmonised consensus (gold standard) from the three different annotators, while also taking concept priority (as specified in the guidelines) into account. We also show that the new Multi-CoreSC CRA corpus allows us to improve performance in the recognition of CoreSCs. The updated guidelines, the multi-label CoreSC CRA corpus and other relevant, related materials are available at the time of publication at http://www.sapientaproject.com/.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesSentence

Similar Papers 제목 키워드 기반

SciCorp: A Corpus of English Scientific Articles Annotated for Information Status Analysis

2016-05-01 · LREC 2016 5 · Ina Roesiger

This paper presents SciCorp, a corpus of full-text English scientific papers of two disciplines, genetics and computational linguistics. The corpus comprises co-reference and bridging information as well as information s…

Articles

Visual Detection with Context for Document Layout Analysis

2019-11-01 · IJCNLP 2019 11 · Carlos Soto, Shinjae Yoo

We present 1) a work in progress method to visually segment key regions of scientific articles using an object detection technique augmented with contextual features, and 2) a novel dataset of region-labeled articles. A …

ArticlesDocument Layout AnalysisLiterature Miningobject-detection+1

Multi-label classification for biomedical literature: an overview of the BioCreative VII LitCovid Track for COVID-19 literature topic annotations

2022-04-20 · Qingyu Chen, Alexis Allot, Robert Leaman, Rezarta Islamaj Doğan 외

The COVID-19 pandemic has been severely impacting global society since December 2019. Massive research has been undertaken to understand the characteristics of the virus and design vaccines and drugs. The related finding…

ArticlesBenchmarkingMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Corpus EN-Istex : un corpus d’articles scientifiques annoté manuellement en entités nommées (ISTEX-EN Corpus: a scientific paper corpus manually annotated in named entities)

2021-06-01 · JEP/TALN/RECITAL 2021 6 · Enza Morale, Denis Maurel, Jeanne Villaneau, Jean-Yves Antoine

Nous présentons ici une nouvelle ressource libre : le corpus EN-ISTEX, un corpus de deux cents articles scientifiques annotés manuellement en entités nommées. Ces articles ont été extraits des deux éditeurs scientifiques…

Articles

SciDTB: Discourse Dependency TreeBank for Scientific Abstracts

2018-06-10 · ACL 2018 7 · An Yang, Sujian Li

Annotation corpus for discourse relations benefits NLP tasks such as machine translation and question answering. In this paper, we present SciDTB, a domain-specific discourse treebank annotated on scientific articles. Di…

ArticlesMachine TranslationQuestion AnsweringTranslation