paper-with-me

C4

Colossal Clean Crawled Corpus

홈페이지 · 논문 981편

C4 is a colossal, cleaned version of Common Crawl's web crawl corpus. It was based on Common Crawl dataset: https://commoncrawl.org. It was used to train the T5 text-to-text Transformer models. The dataset can be downloaded in a pre-processed form from allennlp.

Texts Thai

벤치마크

Language Modelling on C4 결과 9개