Papers Belebele
“Belebele” 태그가 달린 논문 11편 · 필터 해제
Multi-lingual Functional Evaluation for Large Language Models
Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide an adequate understanding of the practi…
BelebeleInstruction FollowingMathMMLUElastic Weight Consolidation for Full-Parameter Continual Pre-Training of Gemma2
This technical report describes an experiment on autoregressive pre-training of Gemma2 2 billion parameter large language model (LLM) with 10\% on the Lithuanian language component of CulturaX from the point of view of c…
ARCBelebeleContinual LearningGSM8K+7Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
Large language models (LLMs) pre-trained predominantly on English text exhibit surprising multilingual capabilities, yet the mechanisms driving cross-lingual generalization remain poorly understood. This work investigate…
BelebeleMachine TranslationNatural Language UnderstandingTranslationDNA 1.0 Technical Report
In this report, we present DNA 1.0 8B Instruct, a state-of-the-art bilingual language model optimized for Korean and English language tasks. By applying continual pre-training (CPT) with high-quality Korean datasets to L…
BelebeleGSM8KInstruction Followingkmmlu+42M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the intersection of BELEBELE and FLEURS, and one s…
BelebeleReading ComprehensionMarco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily English. Many LLMs continue to face challe…
BelebeleMachine TranslationMEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
English-centric large language models (LLMs) often show strong multilingual capabilities. However, the multilingual performance of these models remains unclear and is not thoroughly evaluated for many languages. Most ben…
ARCBelebeleMMLUFrom Multiple-Choice to Extractive QA: A Case Study for English and Arabic
The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data…
BelebeleExtractive Question-AnsweringMachine Reading ComprehensionMultiple-choice+3OpenBA: An Open-sourced 15B Bilingual Asymmetric seq2seq Model Pre-trained from Scratch
Large language models (LLMs) with billions of parameters have demonstrated outstanding performance on various natural language processing tasks. This report presents OpenBA, an open-sourced 15B bilingual asymmetric seq2s…
BelebeleMMLUThe Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. Significantly expanding the language coverage of natural language understanding (NLU) benchmarks, this da…
BelebeleCross-Lingual TransferMachine Reading ComprehensionMultiple-choice+2NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages
In this paper, we create NaijaRC: a new multi-choice Reading Comprehension dataset for three native Nigeria languages that is based on high-school reading comprehension examination. We provide baseline results by perform…
BelebeleCross-Lingual TransferReading Comprehension