paper-with-me

홈 › Papers

C4Corpus: Multilingual Web-size Corpus with Free License

2016-05-01 · LREC 2016 5 · Ivan Habernal, Omnia Zayed, Iryna Gurevych

Large Web corpora containing full documents with permissive licenses are crucial for many NLP tasks. In this article we present the construction of 12 million-pages Web corpus (over 10 billion tokens) licensed under CreativeCommons license family in 50+ languages that has been extracted from CommonCrawl, the largest publicly available general Web crawl to date with about 2 billion crawled URLs. Our highly-scalable Hadoop-based framework is able to process the full CommonCrawl corpus on 2000+ CPU cluster on the Amazon Elastic Map/Reduce infrastructure. The processing pipeline includes license identification, state-of-the-art boilerplate removal, exact duplicate and near-duplicate document removal, and language detection. The construction of the corpus is highly configurable and fully reproducible, and we provide both the framework (DKPro C4CorpusTools) and the resulting data (C4Corpus) to the research community.

📄 PDF Abstract BibTeX

Code (1)

dkpro/dkpro-c4corpus 공식 구현

Tasks

CPU

Similar Papers 제목 키워드 기반

Multilingual Open Text Release 1: Public Domain News in 44 Languages

2022-01-14 · LREC 2022 6 · Chester Palen-Michel, June Kim, Constantine Lignos

We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus cont…

Articles

CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus

2020-02-04 · LREC 2020 5 · Changhan Wang, Juan Pino, Anne Wu, Jiatao Gu

Spoken language translation has recently witnessed a resurgence in popularity, thanks to the development of end-to-end models and the creation of new corpora, such as Augmented LibriSpeech and MuST-C. Existing datasets i…

Speech-to-TextSpeech-to-Text TranslationTranslation

Europarl-ST: A Multilingual Corpus For Speech Translation Of Parliamentary Debates

2019-11-08 · Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Javier Jorge, Nahuel Roselló 외

Current research into spoken language translation (SLT),or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

MultiLegalPile: A 689GB Multilingual Legal Corpus

2023-06-03 · Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis 외

Large, high-quality datasets are crucial for training Large Language Models (LLMs). However, so far, there are few datasets available for specialized critical domains such as law and the available ones are often only for…

Phonetic Segmentation of the UCLA Phonetics Lab Archive

2024-03-28 · Eleanor Chodroff, Blaž Pažon, Annie Baker, Steven Moran

Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio…