paper-with-me

홈 › Papers

A Graph-structured Dataset for Wikipedia Research

2019-03-20 · Nicolas Aspert, Volodymyr Miz, Benjamin Ricaud, Pierre Vandergheynst

Wikipedia is a rich and invaluable source of information. Its central place on the Web makes it a particularly interesting object of study for scientists. Researchers from different domains used various complex datasets related to Wikipedia to study language, social behavior, knowledge organization, and network theory. While being a scientific treasure, the large size of the dataset hinders pre-processing and may be a challenging obstacle for potential new studies. This issue is particularly acute in scientific domains where researchers may not be technically and data processing savvy. On one hand, the size of Wikipedia dumps is large. It makes the parsing and extraction of relevant information cumbersome. On the other hand, the API is straightforward to use but restricted to a relatively small number of requests. The middle ground is at the mesoscopic scale when researchers need a subset of Wikipedia ranging from thousands to hundreds of thousands of pages but there exists no efficient solution at this scale. In this work, we propose an efficient data structure to make requests and access subnetworks of Wikipedia pages and categories. We provide convenient tools for accessing and filtering viewership statistics or "pagecounts" of Wikipedia web pages. The dataset organization leverages principles of graph databases that allows rapid and intuitive access to subgraphs of Wikipedia articles and categories. The dataset and deployment guidelines are available on the LTS2 website \url{https://lts2.epfl.ch/Datasets/Wikipedia/}.

📄 PDF Abstract BibTeX arXiv:1903.08597

Code (1)

epfl-lts2/sparkwiki 공식 구현

Tasks

Articles

Similar Papers 제목 키워드 기반

EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge

2025-07-04 · Klim Zaporojets, Daniel Daza, Edoardo Barba, Ira Assent 외 arxiv

Knowledge Graphs (KGs) are structured knowledge repositories containing entities and relations between them. In this paper, we study the problem of automatically updating KGs over time in response to evolving knowledge i…

Information ExtractionKnowledge Graphs

Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia

2021-01-27 · Tania Chakraborty, Manasa Prasad, Theresa Breiner, Sandy Ritchie 외

Pronunciation modeling is a key task for building speech technology in new languages, and while solid grapheme-to-phoneme (G2P) mapping systems exist, language coverage can stand to be improved. The information needed to…

Biographical: A Semi-Supervised Relation Extraction Dataset

2022-05-02 · Alistair Plum, Tharindu Ranasinghe, Spencer Jones, Constantin Orasan 외

Extracting biographical information from online documents is a popular research topic among the information extraction (IE) community. Various natural language processing (NLP) techniques such as text classification, tex…

ArticlesKnowledge Graphsnamed-entity-recognitionNamed Entity Recognition+6

Classifying Wikipedia in a fine-grained hierarchy: what graphs can contribute

2020-01-21 · Tiphaine Viard, Thomas McLachlan, Hamidreza Ghader, Satoshi Sekine

Wikipedia is a huge opportunity for machine learning, being the largest semi-structured base of knowledge available. Because of this, many works examine its contents, and focus on structuring it in order to make it usabl…

ClassificationGeneral ClassificationSemantic SimilaritySemantic Textual Similarity

WikiGraphs: A Wikipedia Text - Knowledge Graph Paired Dataset

2021-07-20 · NAACL (TextGraphs) 2021 6 · Luyu Wang, Yujia Li, Ozlem Aslan, Oriol Vinyals

We present a new dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning. Existing graph-text paired…

ArticlesConditional Text GenerationGraph GenerationGraph Neural Network+6