C4
Colossal Clean Crawled Corpus
홈페이지 · 논문 981편
C4 is a colossal, cleaned version of Common Crawl's web crawl corpus. It was based on Common Crawl dataset: https://commoncrawl.org. It was used to train the T5 text-to-text Transformer models. The dataset can be downloaded in a pre-processed form from allennlp.
Texts Thai벤치마크
Language Modelling on C4
결과 9개