GerNED: A German Corpus for Named Entity Disambiguation
Determining the real-world referents for name mentions of persons, organizations and other named entities in texts has become an important task in many information retrieval scenarios and is referred to as Named Entity Disambiguation (NED). While comprehensive datasets support the development and evaluation of NED approaches for English, there are no public datasets to assess NED systems for other languages, such as German. This paper describes the construction of an NED dataset based on a large corpus of German news articles. The dataset is closely modeled on the datasets used for the Knowledge Base Population tasks of the Text Analysis Conference, and contains gold standard annotations for the NED tasks of Entity Linking, NIL Detection and NIL Clustering. We also present first experimental results on the new dataset for each of these tasks in order to establish a baseline for future research efforts.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesClusteringCoreference ResolutionEntity DisambiguationEntity LinkingInformation RetrievalKnowledge Base PopulationRetrievalWord Sense DisambiguationSimilar Papers 제목 키워드 기반
Named Entity Corpus Construction using Wikipedia and DBpedia Ontology
In this paper, we propose a novel method to automatically build a named entity corpus based on the DBpedia ontology. Since most of named entity recognition systems require time and effort consuming annotation tasks as tr…
Entity Disambiguationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1A Regional News Corpora for Contextualized Entity Discovery and Linking
This paper presents a German corpus for Named Entity Linking (NEL) and Knowledge Base Population (KBP) tasks. We describe the annotation guideline, the annotation process, NIL clustering techniques and conversion to popu…
ClusteringEntity LinkingKnowledge Base PopulationdiaNED: Time-Aware Named Entity Disambiguation for Diachronic Corpora
Named Entity Disambiguation (NED) systems perform well on news articles and other texts covering a specific time interval. However, NED quality drops when inputs span long time periods like in archives or historic corpor…
ArticlesEntity DisambiguationQTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages
This work presents parallel corpora automatically annotated with several NLP tools, including lemma and part-of-speech tagging, named-entity recognition and classification, named-entity disambiguation, word-sense disambi…
Cross-Lingual TransferEntity DisambiguationGeneral ClassificationLEMMA+7Annotating the MASC Corpus with BabelNet
In this paper we tackle the problem of automatically annotating, with both word senses and named entities, the MASC 3.0 corpus, a large English corpus covering a wide range of genres of written and spoken text. We use Ba…
Entity LinkingReading ComprehensionRelation ExtractionWord Sense Disambiguation