paper-with-me

홈 › Papers

GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages

2024-10-31 · Amir Hossein Kargaran, François Yvon, Hinrich Schütze

The need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models. Most available corpora have sufficient data only for languages with large dominant communities. However, there is no corpus available that (i) covers a wide range of minority languages; (ii) is generated by an open-source reproducible pipeline; and (iii) is rigorously cleaned from noise, making it trustworthy to use. We present GlotCC, a clean, document-level, 2TB general domain corpus derived from CommonCrawl, covering more than 1000 languages. We make GlotCC and the system used to generate it - including the pipeline, language identification model, and filters - available to the research community. Corpus v. 1.0 https://huggingface.co/datasets/cis-lmu/GlotCC-v1, Pipeline v. 3.0 https://github.com/cisnlp/GlotCC.

📄 PDF Abstract BibTeX arXiv:2410.23825

Code (2)

cisnlp/glotcc 공식 구현
🤗 datasets/cis-lmu/GlotCC-V1

Tasks

Language Identification

Similar Papers 제목 키워드 기반

Does Corpus Quality Really Matter for Low-Resource Languages?

2022-03-15 · Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre 외

The vast majority of non-English corpora are derived from automatically filtered versions of CommonCrawl. While prior work has identified major issues on the quality of these datasets (Kreutzer et al., 2021), it is not c…

Representation Learning

FuLG: 150B Romanian Corpus for Language Model Pretraining

2024-07-18 · Vlad-Andrei Bădoiu, Mihai-Valentin Dumitru, Alexandru M. Gherghescu, Alexandru Agache 외

Research in the field of language models is rapidly evolving, with many open models being released to the public. Openly available pretraining corpora usually focus on only a handful of languages, with many others either…

Language ModelingLanguage Modellingmodel

C4Corpus: Multilingual Web-size Corpus with Free License

2016-05-01 · LREC 2016 5 · Ivan Habernal, Omnia Zayed, Iryna Gurevych

Large Web corpora containing full documents with permissive licenses are crucial for many NLP tasks. In this article we present the construction of 12 million-pages Web corpus (over 10 billion tokens) licensed under Crea…

CPU

Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl

2017-10-04 · LREC 2018 5 · Alexander Panchenko, Eugen Ruppert, Stefano Faralli, Simone Paolo Ponzetto 외

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a…

Open Information ExtractionQuestion AnsweringWord Embeddings

A Broad-coverage Corpus for Finnish Named Entity Recognition

2020-05-01 · LREC 2020 5 · Jouni Luoma, Miika Oinonen, Maria Pyyk{\"o}nen, Veronika Laippala 외

We present a new manually annotated corpus for broad-coverage named entity recognition for Finnish. Building on the original Universal Dependencies Finnish corpus of 754 documents (200,000 tokens) representing ten differ…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER