DBpedia Abstracts: A Large-Scale, Open, Multilingual NLP Training Corpus
The ever increasing importance of machine learning in Natural Language Processing is accompanied by an equally increasing need in large-scale training and evaluation corpora. Due to its size, its openness and relative quality, the Wikipedia has already been a source of such data, but on a limited scale. This paper introduces the DBpedia Abstract Corpus, a large-scale, open corpus of annotated Wikipedia texts in six languages, featuring over 11 million texts and over 97 million entity links. The properties of the Wikipedia texts are being described, as well as the corpus creation process, its format and interesting use-cases, like Named Entity Linking training and evaluation.
Code (0)
등록된 구현이 없습니다.
Tasks
Entity LinkingMultilingual NLPSimilar Papers 제목 키워드 기반
DBpedia: A Multilingual Cross-domain Knowledge Base
The DBpedia project extracts structured information from Wikipedia editions in 97 different languages and combines this information into a large multi-lingual knowledge base covering many specific domains and general wor…
Entity LinkingQuestion Answeringslot-fillingSlot Filling+2DBpedia NIF: Open, Large-Scale and Multilingual Knowledge Extraction Corpus
In the past decade, the DBpedia community has put significant amount of effort on developing technical infrastructure and methods for efficient extraction of structured information from Wikipedia. These efforts have been…
ArticlesBuilding Multilingual Corpora for a Complex Named Entity Recognition and Classification Hierarchy using Wikipedia and DBpedia
With the ever-growing popularity of the field of NLP, the demand for datasets in low resourced-languages follows suit. Following a previously established framework, in this paper, we present the UNER dataset, a multiling…
ArticlesNamed Entity RecognitionNamed Entity Recognition (NER)Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation
The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While recent methodologies utilise generic Large Language Models (LLMs) for da…
Data AugmentationMAGES: A Multilingual Angle-integrated Grouping-based Entity Summarization System
This demo presents MAGES (multilingual angle-integrated grouping-based entity summarization), an entity summarization system for a large knowledge base such as DBpedia based on a entity-group-bound ranking in a single in…