paper-with-me

홈 › Papers

A Dataset for Web-Scale Knowledge Base Population

2018-06-03 · European Semantic Web Conference 2018 6 · Michael Glass, Alfio Gliozzo

For many domains, structured knowledge is in short supply, while unstructured text is plentiful. Knowledge Base Population (KBP) is the task of building or extending a knowledge base from text, and systems for KBP have grown in capability and scope. However, existing datasets for KBP are all limited by multiple issues: small in size, not open or accessible, only capable of benchmarking a fraction of the KBP process, or only suitable for extracting knowledge from title-oriented documents (documents that describe a particular entity, such as Wikipedia pages). We introduce and release CC-DBP, a web-scale dataset for training and benchmarking KBP systems. The dataset is based on Common Crawl as the corpus and DBpedia as the target knowledge base. Critically, by releasing the tools to build the dataset, we enable the dataset to remain current as new crawls and DBpedia dumps are released. Also, the modularity of the released tool set resolves a crucial tension between the ease that a dataset can be used for a particular subtask in KBP and the number of different subtasks it can be used to train or benchmark.

📄 PDF Abstract BibTeX

Code (1)

IBM/cc-dbp

Tasks

BenchmarkingKnowledge Base Population

Similar Papers 제목 키워드 기반

Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation Dataset

2021-09-16 · EMNLP 2021 11 · Tianqing Fang, Weiqi Wang, Sehyun Choi, Shibo Hao 외

Reasoning over commonsense knowledge bases (CSKB) whose elements are in the form of free-text is an important yet hard task in NLP. While CSKB completion only fills the missing links within the domain of the CSKB, CSKB p…

BenchmarkingKnowledge Base Population

From Human-Level AI Tales to AI Leveling Human Scales

2026-02-21 · Peter Romero, Fernando Martínez-Plumed, Zachary R. Tidler, Matthieu Téhénan 외 arxiv

Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items again…

PseudoReasoner: Leveraging Pseudo Labels for Commonsense Knowledge Base Population

2022-10-14 · Tianqing Fang, Quyet V. Do, Hongming Zhang, Yangqiu Song 외

Commonsense Knowledge Base (CSKB) Population aims at reasoning over unseen entities and assertions on CSKBs, and is an important yet hard commonsense reasoning task. One challenge is that it requires out-of-domain genera…

Domain GeneralizationKnowledge Base Population

Probabilistic Knowledge Graph Construction: Compositional and Incremental Approaches

2016-08-21 · Dongwoo Kim, Lexing Xie, Cheng Soon Ong

Knowledge graph construction consists of two tasks: extracting information from external resources (knowledge population) and inferring missing information through a statistical analysis on the extracted information (kno…

graph construction

One-shot Transfer Learning for Population Mapping

2021-08-13 · Erzhuo Shao, Jie Feng, Yingheng Wang, Tong Xia 외

Fine-grained population distribution data is of great importance for many applications, e.g., urban planning, traffic scheduling, epidemic modeling, and risk control. However, due to the limitations of data collection, i…

Population MappingSchedulingTransfer Learning