Uncovering the Semantics of Wikipedia Categories
The Wikipedia category graph serves as the taxonomic backbone for large-scale knowledge graphs like YAGO or Probase, and has been used extensively for tasks like entity disambiguation or semantic similarity estimation. Wikipedia's categories are a rich source of taxonomic as well as non-taxonomic information. The category 'German science fiction writers', for example, encodes the type of its resources (Writer), as well as their nationality (German) and genre (Science Fiction). Several approaches in the literature make use of fractions of this encoded information without exploiting its full potential. In this paper, we introduce an approach for the discovery of category axioms that uses information from the category network, category instances, and their lexicalisations. With DBpedia as background knowledge, we discover 703k axioms covering 502k of Wikipedia's categories and populate the DBpedia knowledge graph with additional 4.4M relation assertions and 3.3M type assertions at more than 87% and 90% precision, respectively.
Code (1)
Tasks
Entity DisambiguationKnowledge GraphsSemantic SimilaritySemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Recognizing Descriptive Wikipedia Categories for Historical Figures
Wikipedia is a useful knowledge source that benefits many applications in language processing and knowledge representation. An important feature of Wikipedia is that of categories. Wikipedia pages are assigned different …
DescriptiveInformation RetrievalRetrievalTAGInterpreting Open-Domain Modifiers: Decomposition of Wikipedia Categories into Disambiguated Property-Value Pairs
This paper proposes an open-domain method for automatically annotating modifier constituents (20th-century{'}) within Wikipedia categories (20th-century male writers) with properties (date of birth). The annotations offe…
Mapping WordNet Domains, WordNet Topics and Wikipedia Categories to Generate Multilingual Domain Specific Resources
In this paper we present the mapping between WordNet domains and WordNet topics, and the emergent Wikipedia categories. This mapping leads to a coarse alignment between WordNet and Wikipedia, useful for producing domain-…
Text CategorizationWord Sense DisambiguationCan Wikipedia Categories Improve Masked Language Model Pretraining?
Pretrained language models have obtained impressive results for a large set of natural language understanding tasks. However, training these models is computationally expensive and requires huge amounts of data. Thus, it…
Language ModelingLanguage ModellingNatural Language UnderstandingWhen expertise gone missing: Uncovering the loss of prolific contributors in Wikipedia
Success of planetary-scale online collaborative platforms such as Wikipedia is hinged on active and continued participation of its voluntary contributors. The phenomenal success of Wikipedia as a valued multilingual sour…
Information RetrievalRetrieval