paper-with-me

홈 › Papers

Provenance for Linguistic Corpora Through Nanopublications

2020-06-11 · COLING (LAW) 2020 12 · Timo Lek, Anna de Groot, Tobias Kuhn, Roser Morante

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without significant manual work. Moreover, these annotations might have evolved into different versions, making it challenging for researchers to know the data's provenance. This paper addresses this issue with a case study on event annotated corpora and by creating a new, more interoperable representation of this data in the form of nanopublications. We demonstrate how linguistic annotations from separate corpora can be reliably linked from the start, and thereby be accessed and queried as if they were a single dataset. We describe how such nanopublications can be created and demonstrate how SPARQL queries can be performed to extract interesting content from the new representations. The queries show that information of multiple corpora can be retrieved more easily and effectively because the information of different corpora is represented in a uniform data format.

📄 PDF Abstract BibTeX arXiv:2006.06341

Code (1)

ucds-vu/provcorp-model 공식 구현

Similar Papers 제목 키워드 기반

GAP Enhancing Semantic Interoperability of Genomic Datasets and Provenance Through Nanopublications

2021-11-16 · Matheus Feijoó, Rodrigo Jardim, Sergio Serra, Maria Luiza Campos

While the publication of datasets in scientific repositories has become broadly recognised, the repositories tend to have increasing semantic-related problems. For instance, they present various data reuse obstacles for …

Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl

2017-10-04 · LREC 2018 5 · Alexander Panchenko, Eugen Ruppert, Stefano Faralli, Simone Paolo Ponzetto 외

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a…

Open Information ExtractionQuestion AnsweringWord Embeddings

Human-in-the-Loop Synthetic Text Data Inspection with Provenance Tracking

2024-04-29 · Hong Jin Kang, Fabrice Harel-Canada, Muhammad Ali Gulzar, Violet Peng 외

Data augmentation techniques apply transformations to existing texts to generate additional data. The transformations may produce low-quality texts, where the meaning of the text is changed and the text may even be mangl…

Data AugmentationHate Speech DetectionLanguage ModellingLarge Language Model+1

IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages

2026-03-21 · Sumesh VP arxiv

The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts over two millennia. Despite extensive scholarship on regional Ramay…

Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic Variation

2022-06-01 · LREC 2022 6 · Yaru Wu, Fabian Suchanek, Ioana Vasilescu, Lori Lamel 외

Speech characteristics vary from speaker to speaker. While some variation phenomena are due to the overall communication setting, others are due to diastratic factors such as gender, provenance, age, and social backgroun…