Wikipedia as a Resource for Text Analysis and Retrieval
This tutorial examines the role of Wikipedia in tasks related to text analysis and retrieval. Text analysis tasks, which take advantage of Wikipedia, include coreference resolution, word sense and entity disambiguation and information extraction. In information retrieval, a better understanding of the structure and meaning of queries helps in matching queries against documents, clustering search results, answer and entity retrieval and retrieving knowledge panels for queries asking about popular entities.
Code (0)
등록된 구현이 없습니다.
Tasks
Clusteringcoreference-resolutionCoreference ResolutionEntity DisambiguationEntity RetrievalInformation RetrievalRetrievalSimilar Papers 제목 키워드 기반
The Role of Wikipedia in Text Analysis and Retrieval
This tutorial examines the characteristics, advantages and limitations of Wikipedia relative to other existing, human-curated resources of knowledge; derivative resources, created by converting semi-structured content in…
Coreference ResolutionInformation RetrievalRetrievalWikipedia Text Reuse: Within and Without
We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse b…
RetrievalDetecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is the…
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
In this paper, we present a novel series of Russian information retrieval datasets constructed from the "Did you know..." section of Russian Wikipedia. Our datasets support a range of retrieval tasks, including fact-chec…
Information RetrievalDesign Challenges in Low-resource Cross-lingual Entity Linking
Cross-lingual Entity Linking (XEL), the problem of grounding mentions of entities in a foreign language text into an English knowledge base such as Wikipedia, has seen a lot of research in recent years, with a range of p…
Cross-Lingual Entity LinkingEntity Linking