paper-with-me

Papers

How GermaParl Evolves: Improving Data Quality by Reproducible Corpus Preparation and User Involvement

2022-06-01 · ParlaCLARIN (LREC) 2022 6 · Andreas Blaette, Julia Rakers, Christoph Leonhardt

The development and curation of large-scale corpora of plenary debates requires not only care and attention to detail when the data is created but also effective means of sustainable quality control. This paper makes two contributions: Firstly, it presents an updated version of the GermaParl corpus of parliamentary debates in the German *Bundestag*. Secondly, it shows how the corpus preparation pipeline is designed to serve the quality of the resource by facilitating effective community involvement. Centered around a workflow which combines reproducibility, transparency and version control, the pipeline allows for continuous improvements to the corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The GermaParl Corpus of Parliamentary Protocols

2018-05-01 · LREC 2018 5 · Andreas Bl{\"a}tte, Andre Blessing
Decision MakingMachine Translation

Lit2Vec: A Reproducible Workflow for Building a Legally Screened Chemistry Corpus from S2ORC for Downstream Retrieval and Text Mining

2026-04-14 · Mahmoud Amiri, Jamile Mohammad Jafari, Sara Mostafapour, Thomas Bocklitz arxiv

We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we …

PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development

2026-03-17 · Hanif Rahman arxiv

We present PashtoCorp, a 1.25-billion-word corpus for Pashto, a language spoken by 60 million people that remains severely underrepresented in NLP. The corpus is assembled from 39 sources spanning seven HuggingFace datas…

Reading Comprehension

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

2025-01-14 · Yijiong Yu, Ziyun Dai, Zekun Wang, Wei Wang 외

Large language models (LLMs) have demonstrated remarkable capabilities, but their success heavily relies on the quality of pretraining corpora. For Chinese LLMs, the scarcity of high-quality Chinese datasets presents a s…

Stem-cell differentiation underpins reproducible morphogenesis

2025-03-25 · Dominic K Devlin, Austen RD Ganley, Nobuto Takeuchi

Morphogenesis of complex body shapes is reproducible despite the noise inherent in the underlying morphogenetic processes. However, how these morphogenetic processes work together to achieve this reproducibility remains …