paper-with-me

홈 › Papers

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

2021-04-18 · EMNLP 2021 11 · Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner

Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentation. In this work we provide some of the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin by investigating where the data came from, and find a significant amount of text from unexpected sources like patents and US military websites. Then we explore the content of the text itself, and find machine-generated text (e.g., from machine translation systems) and evaluation examples from other benchmark NLP datasets. To understand the impact of the filters applied to create this dataset, we evaluate the text that was removed, and show that blocklist filtering disproportionately removes text from and about minority individuals. Finally, we conclude with some recommendations for how to created and document web-scale datasets from a scrape of the internet.

📄 PDF Abstract BibTeX arXiv:2104.08758

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Context-sensitive evaluation of automatic speech recognition: considering user experience & language variation

2021-04-01 · EACL (HCINLP) 2021 4 · Nina Markl, Catherine Lai

Commercial Automatic Speech Recognition (ASR) systems tend to show systemic predictive bias for marginalised speaker/user groups. We highlight the need for an interdisciplinary and context-sensitive approach to documenti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Know thy corpus! Robust methods for digital curation of Web corpora

2020-03-13 · LREC 2020 5 · Serge Sharoff

This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained …

Genre classificationTopic Models

ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information

2024-11-29 · Wanyue Zhang, Ziyong Li, Wen Yang, Chunlin Leng 외

During the development of large language models (LLMs), pre-training data play a critical role in shaping LLMs' capabilities. In recent years several large-scale and high-quality pre-training datasets have been released …

WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset

2024-02-29 · Jiantao Qiu, Haijun Lv, Zhenjiang Jin, Rui Wang 외

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing large-scale pre-training datasets for langua…

Bad Seeds: Evaluating Lexical Methods for Bias Measurement

2021-08-01 · ACL 2021 5 · Maria Antoniak, David Mimno

A common factor in bias measurement methods is the use of hand-curated seed lexicons, but there remains little guidance for their selection. We gather seeds used in prior work, documenting their common sources and ration…