paper-with-me

Liu et al. Corpus

· 논문 1편

The Liu et al. Corpus is a pretraining dataset for large language models. It consists of 160Gb of news, books, stories, and web text.

Texts English