paper-with-me

홈 › Papers

BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models

2025-02-11 · Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei LI, ShuJian Huang, Fei Yuan

Previous multilingual benchmarks focus primarily on simple understanding tasks, but for large language models(LLMs), we emphasize proficiency in instruction following, reasoning, long context understanding, code generation, and so on. However, measuring these advanced capabilities across languages is underexplored. To address the disparity, we introduce BenchMAX, a multi-way multilingual evaluation benchmark that allows for fair comparisons of these important abilities across languages. To maintain high quality, three distinct native-speaking annotators independently annotate each sample within all tasks after the data was machine-translated from English into 16 other languages. Additionally, we present a novel translation challenge stemming from dataset construction. Extensive experiments on BenchMAX reveal varying effectiveness of core capabilities across languages, highlighting performance gaps that cannot be bridged by simply scaling up model size. BenchMAX serves as a comprehensive multilingual evaluation platform, providing a promising test bed to promote the development of multilingual language models. The dataset and code are publicly accessible.

📄 PDF Abstract BibTeX arXiv:2502.07346

Code (1)

cone-mt/benchmax 공식 구현

Tasks

Code GenerationInstruction FollowingLong-Context Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MTG: A Benchmark Suite for Multilingual Text Generation

2021-08-13 · Findings (NAACL) 2022 7 · Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang 외

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first-proposed multilingual multiway text generation dataset with the largest human-annotated data (400k). It in…

Question GenerationQuestion-GenerationStory GenerationText Generation+2

LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models

2025-08-18 · Zhiyuan Ning, Tianle Gu, Jiaxin Song, Shixin Hong 외 arxiv

The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse range of linguistic and cultural contexts. T…

SEA-HELM: Southeast Asian Holistic Evaluation of Language Models

2025-02-20 · Yosephine Susanto, Adithya Venkatadri Hulagadri, Jann Railey Montalan, Jian Gang Ngui 외

With the rapid emergence of novel capabilities in Large Language Models (LLMs), the need for rigorous multilingual and multicultural benchmarks that are integrated has become more pronounced. Though existing LLM benchmar…

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

2025-05-31 · Shaoxiong Ji, Zihao Li, Jaakko Paavola, Indraneil Paul 외

This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for m…

Translation

Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models

2025-07-02 · Samridhi Raj Sinha, Rajvee Sheth, Abhishek Upperwal, Mayank Singh arxiv

The rapid evolution of Large Language Models' has underscored the need for evaluation frameworks that are globally applicable, flexible, and modular, and that support a wide range of tasks, model types, and linguistic se…