paper-with-me

Papers

DBpedia NIF: Open, Large-Scale and Multilingual Knowledge Extraction Corpus

2018-12-26 · Milan Dojchinovski, Julio Hernandez, Markus Ackermann, Amit Kirschenbaum, Sebastian Hellmann

In the past decade, the DBpedia community has put significant amount of effort on developing technical infrastructure and methods for efficient extraction of structured information from Wikipedia. These efforts have been primarily focused on harvesting, refinement and publishing semi-structured information found in Wikipedia articles, such as information from infoboxes, categorization information, images, wikilinks and citations. Nevertheless, still vast amount of valuable information is contained in the unstructured Wikipedia article texts. In this paper, we present DBpedia NIF - a large-scale and multilingual knowledge extraction corpus. The aim of the dataset is two-fold: to dramatically broaden and deepen the amount of structured information in DBpedia, and to provide large-scale and multilingual language resource for development of various NLP and IR task. The dataset provides the content of all articles for 128 Wikipedia languages. We describe the dataset creation process and the NLP Interchange Format (NIF) used to model the content, links and the structure the information of the Wikipedia articles. The dataset has been further enriched with about 25% more links and selected partitions published as Linked Data. Finally, we describe the maintenance and sustainability plans, and selected use cases of the dataset from the TextExt knowledge extraction challenge.

📄 PDF Abstract BibTeX arXiv:1812.10315

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

DBpedia Abstracts: A Large-Scale, Open, Multilingual NLP Training Corpus

2016-05-01 · LREC 2016 5 · Martin Br{\"u}mmer, Milan Dojchinovski, Sebastian Hellmann

The ever increasing importance of machine learning in Natural Language Processing is accompanied by an equally increasing need in large-scale training and evaluation corpora. Due to its size, its openness and relative qu…

Entity LinkingMultilingual NLP

DBpedia: A Multilingual Cross-domain Knowledge Base

2012-05-01 · LREC 2012 5 · Pablo Mendes, Max Jakob, Christian Bizer

The DBpedia project extracts structured information from Wikipedia editions in 97 different languages and combines this information into a large multi-lingual knowledge base covering many specific domains and general wor…

Entity LinkingQuestion Answeringslot-fillingSlot Filling+2

MAGES: A Multilingual Angle-integrated Grouping-based Entity Summarization System

2016-12-01 · COLING 2016 12 · Eun-Kyung Kim, Key-Sun Choi

This demo presents MAGES (multilingual angle-integrated grouping-based entity summarization), an entity summarization system for a large knowledge base such as DBpedia based on a entity-group-bound ranking in a single in…

On-Demand and Lightweight Knowledge Graph Generation -- a Demonstration with DBpedia

2021-07-02 · Malte Brockmeier, Yawen Liu, Sunita Pateer, Sven Hertling 외

Modern large-scale knowledge graphs, such as DBpedia, are datasets which require large computational resources to serve and process. Moreover, they often have longer release cycles, which leads to outdated information in…

Graph GenerationKnowledge Graphs

EventKG: A Multilingual Event-Centric Temporal Knowledge Graph

2018-04-12 · Simon Gottschalk, Elena Demidova

One of the key requirements to facilitate semantic analytics of information regarding contemporary and historical events on the Web, in the news and in social media is the availability of reference knowledge repositories…

Knowledge Graphs