paper-with-me

Papers Belebele

“Belebele” 태그가 달린 논문 11편 · 필터 해제

Multi-lingual Functional Evaluation for Large Language Models

2025-06-25 · Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian

Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide an adequate understanding of the practi…

BelebeleInstruction FollowingMathMMLU

Elastic Weight Consolidation for Full-Parameter Continual Pre-Training of Gemma2

2025-05-09 · Vytenis Šliogeris, Povilas Daniušis, Artūras Nakvosas

This technical report describes an experiment on autoregressive pre-training of Gemma2 2 billion parameter large language model (LLM) with 10\% on the Lithuanian language component of CulturaX from the point of view of c…

ARCBelebeleContinual LearningGSM8K+7

Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs

2025-04-13 · Kartik Ravisankar, Hyojung Han, Marine Carpuat

Large language models (LLMs) pre-trained predominantly on English text exhibit surprising multilingual capabilities, yet the mechanisms driving cross-lingual generalization remain poorly understood. This work investigate…

BelebeleMachine TranslationNatural Language UnderstandingTranslation

DNA 1.0 Technical Report

2025-01-18 · Jungyup Lee, Jemin Kim, Sang Park, Seungjae Lee

In this report, we present DNA 1.0 8B Instruct, a state-of-the-art bilingual language model optimized for Korean and English language tasks. By applying continual pre-training (CPT) with high-quality Korean datasets to L…

BelebeleGSM8KInstruction Followingkmmlu+4

2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset

2024-12-11 · Marta R. Costa-jussà, Bokai Yu, Pierre Andrews, Belen Alastruey 외

We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the intersection of BELEBELE and FLEURS, and one s…

BelebeleReading Comprehension

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement

2024-12-05 · Lingfeng Ming, Bo Zeng, Chenyang Lyu, Tianqi Shi 외

Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily English. Many LLMs continue to face challe…

BelebeleMachine Translation

MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment

2024-10-08 · Amir Hossein Kargaran, Ali Modarressi, Nafiseh Nikeghbal, Jana Diesner 외

English-centric large language models (LLMs) often show strong multilingual capabilities. However, the multilingual performance of these models remains unclear and is not thoroughly evaluated for many languages. Most ben…

ARCBelebeleMMLU

From Multiple-Choice to Extractive QA: A Case Study for English and Arabic

2024-04-26 · Teresa Lynn, Malik H. Altakrori, Samar Mohamed Magdy, Rocktim Jyoti Das 외

The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data…

BelebeleExtractive Question-AnsweringMachine Reading ComprehensionMultiple-choice+3

OpenBA: An Open-sourced 15B Bilingual Asymmetric seq2seq Model Pre-trained from Scratch

2023-09-19 · Juntao Li, Zecheng Tang, Yuyang Ding, Pinzheng Wang 외

Large language models (LLMs) with billions of parameters have demonstrated outstanding performance on various natural language processing tasks. This report presents OpenBA, an open-sourced 15B bilingual asymmetric seq2s…

BelebeleMMLU

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

2023-08-31 · Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 외

We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. Significantly expanding the language coverage of natural language understanding (NLU) benchmarks, this da…

BelebeleCross-Lingual TransferMachine Reading ComprehensionMultiple-choice+2

NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages

2023-08-18 · Anuoluwapo Aremu, Jesujoba O. Alabi, Daud Abolade, Nkechinyere F. Aguobi 외

In this paper, we create NaijaRC: a new multi-choice Reading Comprehension dataset for three native Nigeria languages that is based on high-school reading comprehension examination. We provide baseline results by perform…

BelebeleCross-Lingual TransferReading Comprehension
1–11 / 11