paper-with-me

홈 › Papers

Wiki-40B: Multilingual Language Model Dataset

2020-05-01 · LREC 2020 5 · M. Guo, y, Zihang Dai, Vr, Denny e{\v{c}}i{\'c}, Rami Al-Rfou

We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the research of multilingual modeling. We train monolingual causal language models using a state-of-the-art model (Transformer-XL) establishing baselines for many languages. We also introduce the task of multilingual causal language modeling where we train our model on the combined text of 40+ languages from Wikipedia with different vocabulary sizes and evaluate on the languages individually. We released the cleaned-up text of 40+ Wikipedia language editions, the corresponding trained monolingual language models, and several multilingual language models with different fixed vocabulary sizes.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Language ModelingLanguage ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

Considerations for Multilingual Wikipedia Research

2022-04-05 · Isaac Johnson, Emily Lescak

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and ca…

Factual Inconsistencies in Multilingual Wikipedia Tables

2025-07-24 · Silvia Cappa, Lingxiao Kong, Pille-Riin Peet, Fanfu Wei 외 arxiv

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to fa…

Understanding Editing Behaviors in Multilingual Wikipedia

2015-08-28 · Suin Kim, Sungjoon Park, Scott A. Hale, Sooyoung Kim 외

Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different lan…

An Open Multilingual System for Scoring Readability of Wikipedia

2024-06-03 · Mykola Trokhymovych, Indira Sen, Martin Gerlach

With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the…

Articles

X-WikiRE: A Large, Multilingual Resource for Relation Extraction as Machine Comprehension

2019-08-14 · WS 2019 11 · Mostafa Abdou, Cezar Sas, Rahul Aralikatte, Isabelle Augenstein 외

Although the vast majority of knowledge bases KBs are heavily biased towards English, Wikipedias do cover very different topics in different languages. Exploiting this, we introduce a new multilingual dataset (X-WikiRE),…

Reading ComprehensionRelationRelation Extraction