Hedera: Scalable Indexing and Exploring Entities in Wikipedia Revision History
Much of work in semantic web relying on Wikipedia as the main source of knowledge often work on static snapshots of the dataset. The full history of Wikipedia revisions, while contains much more useful information, is still difficult to access due to its exceptional volume. To enable further research on this collection, we developed a tool, named Hedera, that efficiently extracts semantic information from Wikipedia revision history datasets. Hedera exploits Map-Reduce paradigm to achieve rapid extraction, it is able to handle one entire Wikipedia articles revision history within a day in a medium-scale cluster, and supports flexible data structures for various kinds of semantic web study.
Code (1)
Tasks
ArticlesSimilar Papers 제목 키워드 기반
RELink: A Research Framework and Test Collection for Entity-Relationship Retrieval
Improvements of entity-relationship (E-R) search techniques have been hampered by a lack of test collections, particularly for complex queries involving multiple entities and relationships. In this paper we describe a me…
Natural Language QueriesRetrievalExploring semantically-related concepts from Wikipedia: the case of SeRE
In this paper we present our web application SeRE designed to explore semantically related concepts. Wikipedia and DBpedia are rich data sources to extract related entities for a given topic, like in- and out-links, broa…
ArticlesGeneral ClassificationEntity Cloze By Date: What LMs Know About Unseen Entities
Language models (LMs) are typically trained once on a large-scale corpus and used for years without being updated. However, in a dynamic world, new entities constantly arise. We propose a framework to analyze what LMs ca…
ArticlesDyVo: Dynamic Vocabularies for Learned Sparse Retrieval with Entities
Learned Sparse Retrieval (LSR) models use vocabularies from pre-trained transformers, which often split entities into nonsensical fragments. Splitting entities can reduce retrieval accuracy and limits the model's ability…
Document RankingEntity EmbeddingsEntity RetrievalRetrieval+1Entity Cloze By Date: Understanding what LMs know about unseen entities
Language models (LMs) are typically trained once on a large-scale corpus and used for years without being updated. Our world, however, is dynamic, and new entities constantly arise. We propose a framework to analyze what…
ArticlesDate Understanding