paper-with-me

홈 › Papers

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

2025-03-13 · Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Yun Xing, Junjue Wang, Huitao Li, Xin Li, Kunyu Yu, Nan Liu, Qingyu Chen, Douglas Teodoro, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li

Traditional benchmarks struggle to evaluate increasingly sophisticated language models in multilingual and culturally diverse contexts. To address this gap, we introduce MMLU-ProX, a comprehensive multilingual benchmark covering 13 typologically diverse languages with approximately 11,829 questions per language. Building on the challenging reasoning-focused design of MMLU-Pro, our framework employs a semi-automatic translation process: translations generated by state-of-the-art large language models (LLMs) are rigorously evaluated by expert annotators to ensure conceptual accuracy, terminological consistency, and cultural relevance. We comprehensively evaluate 25 state-of-the-art LLMs using 5-shot chain-of-thought (CoT) and zero-shot prompting strategies, analyzing their performance across linguistic and cultural boundaries. Our experiments reveal consistent performance degradation from high-resource languages to lower-resource ones, with the best models achieving over 70% accuracy on English but dropping to around 40% for languages like Swahili, highlighting persistent gaps in multilingual capabilities despite recent advances. MMLU-ProX is an ongoing project; we are expanding our benchmark by incorporating additional languages and evaluating more language models to provide a more comprehensive assessment of multilingual capabilities.

📄 PDF Abstract BibTeX arXiv:2503.10497

Code (0)

등록된 구현이 없습니다.

Tasks

Language Model EvaluationLanguage ModelingLanguage ModellingLarge Language ModelMMLU

Similar Papers 제목 키워드 기반

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…

CMMLU: Measuring massive multitask language understanding in Chinese

2023-06-15 · Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang 외

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance becomes increasingly crucial and challenging. This paper aims to bridge this gap by introducing CMMLU, a comprehensive…

Large Language Model

Measuring Hong Kong Massive Multi-Task Language Understanding

2025-05-04 · Chuxue Cao, Zhenghao Zhu, Junqi Zhu, Guoying Lu 외

Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional …

MMLUMulti-task Language Understanding

TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages

2025-02-16 · Jafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova 외

Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However, preparing such benchmarks in high quali…

Machine TranslationMMLUMulti-task Language Understanding

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

2024-12-04 · Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani 외

Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from language but also from the cultural knowledge required to interpret quest…

MMLU