Entity Extraction from Wikipedia List Pages
When it comes to factual knowledge about a wide range of domains, Wikipedia is often the prime source of information on the web. DBpedia and YAGO, as large cross-domain knowledge graphs, encode a subset of that knowledge by creating an entity for each page in Wikipedia, and connecting them through edges. It is well known, however, that Wikipedia-based knowledge graphs are far from complete. Especially, as Wikipedia's policies permit pages about subjects only if they have a certain popularity, such graphs tend to lack information about less well-known entities. Information about these entities is oftentimes available in the encyclopedia, but not represented as an individual page. In this paper, we present a two-phased approach for the extraction of entities from Wikipedia's list pages, which have proven to serve as a valuable source of information. In the first phase, we build a large taxonomy from categories and list pages with DBpedia as a backbone. With distant supervision, we extract training data for the identification of new entities in list pages that we use in the second phase to train a classification model. With this approach we extract over 700k new entities and extend DBpedia with 7.5M new type statements and 3.8M new facts of high precision.
Code (0)
등록된 구현이 없습니다.
Tasks
Entity Extraction using GANKnowledge GraphsSimilar Papers 제목 키워드 기반
Automated News Suggestions for Populating Wikipedia Entity Pages
Wikipedia entity pages are a valuable source of information for direct consumption and for knowledge-base construction, update and maintenance. Facts in these entity pages are typically supported by references. Recent st…
ArticlesKnowledge Base ConstructionEntity Linking with people entity on Wikipedia
This paper introduces a new model that uses named entity recognition, coreference resolution, and entity linking techniques, to approach the task of linking people entities on Wikipedia people pages to their correspondin…
coreference-resolutionCoreference ResolutionEntity Linkingnamed-entity-recognition+2Resource of Wikipedias in 31 Languages Categorized into Fine-Grained Named Entities
This paper describes a resource of Wikipedias in 31 languages categorized into Extended Named Entity (ENE), which has 219 fine-grained NE categories. We first categorized 920 K Japanese Wikipedia pages according to the E…
AttributeAttribute ExtractionEnsemble LearningLink PredictionEntity Disambiguation with Web Links
Entity disambiguation with Wikipedia relies on structured information from redirect pages, article text, inter-article links, and categories. We explore whether web links can replace a curated encyclopaedia, obtaining en…
Entity DisambiguationEntity LinkingKnowledge Base PopulationDistant Supervision for Entity Linking
Entity linking is an indispensable operation of populating knowledge repositories for information extraction. It studies on aligning a textual entity mention to its corresponding disambiguated entry in a knowledge reposi…
DescriptiveEntity Linking