paper-with-me

홈 › Papers

Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia

2025-03-18 · Jan Göpfert, Patrick Kuckertz, Jann M. Weinand, Detlef Stolten

To cope with the large number of publications, more and more researchers are automatically extracting data of interest using natural language processing methods based on supervised learning. Much data, especially in the natural and engineering sciences, is quantitative, but there is a lack of datasets for identifying quantities and their context in text. To address this issue, we present two large datasets based on Wikipedia and Wikidata: Wiki-Quantities is a dataset consisting of over 1.2 million annotated quantities in the English-language Wikipedia. Wiki-Measurements is a dataset of 38,738 annotated quantities in the English-language Wikipedia along with their respective measured entity, property, and optional qualifiers. Manual validation of 100 samples each of Wiki-Quantities and Wiki-Measurements found 100% and 84-94% correct, respectively. The datasets can be used in pipeline approaches to measurement extraction, where quantities are first identified and then their measurement context. To allow reproduction of this work using newer or different versions of Wikipedia and Wikidata, we publish the code used to create the datasets along with the data.

📄 PDF Abstract BibTeX arXiv:2503.14090

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mining Knowledge for Natural Language Inference from Wikipedia Categories

2020-10-03 · Findings of the Association for Computational Linguistics 2020 · Mingda Chen, Zewei Chu, Karl Stratos, Kevin Gimpel

Accurate lexical entailment (LE) and natural language inference (NLI) often require large quantities of costly annotations. To alleviate the need for labeled data, we introduce WikiNLI: a resource for improving model per…

Lexical EntailmentNatural Language Inference

Mining Naturally-occurring Corrections and Paraphrases from Wikipedia's Revision History

2022-02-25 · Aurélien Max, Guillaume Wisniewski

Naturally-occurring instances of linguistic phenomena are important both for training and for evaluating automatic processes on text. When available in large quantities, they also prove interesting material for linguisti…

WIKIPARQ: A Tabulated Wikipedia Resource Using the Parquet Format

2016-05-01 · LREC 2016 5 · Marcus Klang, Pierre Nugues

Wikipedia has become one of the most popular resources in natural language processing and it is used in quantities of applications. However, Wikipedia requires a substantial pre-processing step before it can be used. For…

AllArticles

TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents

2026-08-02 · Jingyu Sun, Yuyang Xue, Mingyang Li, Zhengtao Yao 외 arxiv

Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable,…

Answer Generation

INRIASAC: Simple Hypernym Extraction Methods

2015-02-04 · SEMEVAL 2015 6 · Gregory Grefenstette

Given a set of terms from a given domain, how can we structure them into a taxonomy without manual intervention? This is the task 17 of SemEval 2015. Here we present our simple taxonomy structuring techniques which, desp…

Sentence