Corpus for Coreference Resolution on Scientific Papers
The ever-growing number of published scientific papers prompts the need for automatic knowledge extraction to help scientists keep up with the state-of-the-art in their respective fields. To construct a good knowledge extraction system, annotated corpora in the scientific domain are required to train machine learning models. As described in this paper, we have constructed an annotated corpus for coreference resolution in multiple scientific domains, based on an existing corpus. We have modified the annotation scheme from Message Understanding Conference to better suit scientific texts. Then we applied that to the corpus. The annotated corpus is then compared with corpora in general domains in terms of distribution of resolution classes and performance of the Stanford Dcoref coreference resolver. Through these comparisons, we have demonstrated quantitatively that our manually annotated corpus differs from a general-domain corpus, which suggests deep differences between general-domain texts and scientific texts and which shows that different approaches can be made to tackle coreference resolution for general texts and scientific texts.
Code (1)
Tasks
coreference-resolutionCoreference ResolutionOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
Coreference Resolution in Research Papers from Multiple Domains
Coreference resolution is essential for automatic text understanding to facilitate high-level information retrieval tasks such as text summarisation or question answering. Previous work indicates that the performance of …
coreference-resolutionCoreference ResolutionInformation RetrievalRetrieval+1SciCo: Hierarchical Cross-Document Coreference for Scientific Concepts
Determining coreference of concept mentions across multiple documents is a fundamental task in natural language understanding. Previous work on cross-document coreference resolution (CDCR) typically considers mentions of…
coreference-resolutionCoreference ResolutionCross Document Coreference ResolutionNatural Language UnderstandingMarmara Turkish Coreference Corpus and Coreference Resolution Baseline
We describe the Marmara Turkish Coreference Corpus, which is an annotation of the whole METU-Sabanci Turkish Treebank with mentions and coreference chains. Collecting eight or more independent annotations for each docume…
coreference-resolutionCoreference ResolutionqxoRef 1.0: A coreference corpus and mention-pair baseline for coreference resolution in Conchucos Quechua
This paper introduces qxoRef 1.0, the first coreference corpus to be developed for a Quechuan language, and describes a baseline mention-pair coreference resolution system developed for this corpus. The evaluation of thi…
coreference-resolutionCoreference ResolutionSciCorp: A Corpus of English Scientific Articles Annotated for Information Status Analysis
This paper presents SciCorp, a corpus of full-text English scientific papers of two disciplines, genetics and computational linguistics. The corpus comprises co-reference and bridging information as well as information s…
Articles