paper-with-me

홈 › Papers

WIKIPARQ: A Tabulated Wikipedia Resource Using the Parquet Format

2016-05-01 · LREC 2016 5 · Marcus Klang, Pierre Nugues

Wikipedia has become one of the most popular resources in natural language processing and it is used in quantities of applications. However, Wikipedia requires a substantial pre-processing step before it can be used. For instance, its set of nonstandardized annotations, referred to as the wiki markup, is language-dependent and needs specific parsers from language to language, for English, French, Italian, etc. In addition, the intricacies of the different Wikipedia resources: main article text, categories, wikidata, infoboxes, scattered into the article document or in different files make it difficult to have global view of this outstanding resource. In this paper, we describe WikiParq, a unified format based on the Parquet standard to tabulate and package the Wikipedia corpora. In combination with Spark, a map-reduce computing framework, and the SQL query language, WikiParq makes it much easier to write database queries to extract specific information or subcorpora from Wikipedia, such as all the first paragraphs of the articles in French, or all the articles on persons in Spanish, or all the articles on persons that have versions in French, English, and Spanish. WikiParq is available in six language versions and is potentially extendible to all the languages of Wikipedia. The WikiParq files are downloadable as tarball archives from this location: http://semantica.cs.lth.se/wikiparq/.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AllArticles

Similar Papers 제목 키워드 기반

TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification

2026-05-09 · Xinyu Chen, HanQin Cai, Lijun Ding, Jinhua Zhao arxiv

We present TailedTS, a large-scale benchmark dataset derived from Wikipedia hourly page view observations throughout 2024, specifically designed to test time series forecasting models under heavy-tailed, zero-inflated, a…

Time Series ForecastingTime Series Prediction

Tuning Parameter-Free Nonparametric Density Estimation from Tabulated Summary Data

2022-04-12 · Ji Hyung Lee, Yuya Sasaki, Alexis Akira Toda, Yulong Wang

Administrative data are often easier to access as tabulated summaries than in the original format due to confidentiality concerns. Motivated by this practical feature, we propose a novel nonparametric density estimation …

Density Estimation

INFOTABS: Inference on Tables as Semi-structured Data

2020-05-13 · ACL 2020 6 · Vivek Gupta, Maitrey Mehta, Pegah Nokhiz, Vivek Srikumar

In this paper, we observe that semi-structured tabulated text is ubiquitous; understanding them requires not only comprehending the meaning of text fragments, but also implicit relationships between them. We argue that s…

The Role of Wikipedia in Text Analysis and Retrieval

2016-12-01 · COLING 2016 12 · Marius Pa{\c{s}}ca

This tutorial examines the characteristics, advantages and limitations of Wikipedia relative to other existing, human-curated resources of knowledge; derivative resources, created by converting semi-structured content in…

Coreference ResolutionInformation RetrievalRetrieval

Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia

2021-01-27 · Tania Chakraborty, Manasa Prasad, Theresa Breiner, Sandy Ritchie 외

Pronunciation modeling is a key task for building speech technology in new languages, and while solid grapheme-to-phoneme (G2P) mapping systems exist, language coverage can stand to be improved. The information needed to…