paper-with-me

Papers

Hedera: Scalable Indexing and Exploring Entities in Wikipedia Revision History

2017-01-14 · Tuan Tran, Tu Ngoc Nguyen

Much of work in semantic web relying on Wikipedia as the main source of knowledge often work on static snapshots of the dataset. The full history of Wikipedia revisions, while contains much more useful information, is still difficult to access due to its exceptional volume. To enable further research on this collection, we developed a tool, named Hedera, that efficiently extracts semantic information from Wikipedia revision history datasets. Hedera exploits Map-Reduce paradigm to achieve rapid extraction, it is able to handle one entire Wikipedia articles revision history within a day in a medium-scale cluster, and supports flexible data structures for various kinds of semantic web study.

📄 PDF Abstract BibTeX arXiv:1701.03937

Code (1)

antoine-tran/Hedera 공식 구현

Tasks

Articles

Similar Papers 제목 키워드 기반

RELink: A Research Framework and Test Collection for Entity-Relationship Retrieval

2017-06-13 · Saleiro Pedro, Milic-Frayling Natasa, Rodrigues Eduarda Mendes, Soares Carlos

Improvements of entity-relationship (E-R) search techniques have been hampered by a lack of test collections, particularly for complex queries involving multiple entities and relationships. In this paper we describe a me…

Natural Language QueriesRetrieval

Exploring semantically-related concepts from Wikipedia: the case of SeRE

2015-04-27 · Daniel Hienert, Dennis Wegener, Siegfried Schomisch

In this paper we present our web application SeRE designed to explore semantically related concepts. Wikipedia and DBpedia are rich data sources to extract related entities for a given topic, like in- and out-links, broa…

ArticlesGeneral Classification

Entity Cloze By Date: What LMs Know About Unseen Entities

2022-05-05 · Findings (NAACL) 2022 7 · Yasumasa Onoe, Michael J. Q. Zhang, Eunsol Choi, Greg Durrett

Language models (LMs) are typically trained once on a large-scale corpus and used for years without being updated. However, in a dynamic world, new entities constantly arise. We propose a framework to analyze what LMs ca…

Articles

DyVo: Dynamic Vocabularies for Learned Sparse Retrieval with Entities

2024-10-10 · Thong Nguyen, Shubham Chatterjee, Sean MacAvaney, Iain Mackie 외

Learned Sparse Retrieval (LSR) models use vocabularies from pre-trained transformers, which often split entities into nonsensical fragments. Splitting entities can reduce retrieval accuracy and limits the model's ability…

Document RankingEntity EmbeddingsEntity RetrievalRetrieval+1

Entity Cloze By Date: Understanding what LMs know about unseen entities

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Language models (LMs) are typically trained once on a large-scale corpus and used for years without being updated. Our world, however, is dynamic, and new entities constantly arise. We propose a framework to analyze what…

ArticlesDate Understanding