paper-with-me

LTI LangID Corpus

홈페이지 · 논문 1편

The LTI LangID Corpus is a dataset used for language identification (LangID) tasks. It contains text data in various languages. The dataset has had multiple releases, with the first release containing 781 "core" languages and 1091 languages overall.