paper-with-me

Papers

CMMLU: Measuring massive multitask language understanding in Chinese

2023-06-15 · Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, Timothy Baldwin

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance becomes increasingly crucial and challenging. This paper aims to bridge this gap by introducing CMMLU, a comprehensive Chinese benchmark that covers various subjects, including natural science, social sciences, engineering, and humanities. We conduct a thorough evaluation of 18 advanced multilingual- and Chinese-oriented LLMs, assessing their performance across different subjects and settings. The results reveal that most existing LLMs struggle to achieve an average accuracy of 50%, even when provided with in-context examples and chain-of-thought prompts, whereas the random baseline stands at 25%. This highlights significant room for improvement in LLMs. Additionally, we conduct extensive experiments to identify factors impacting the models' performance and propose directions for enhancing LLMs. CMMLU fills the gap in evaluating the knowledge and reasoning capabilities of large language models within the Chinese context.

📄 PDF Abstract BibTeX arXiv:2306.09212

Code (1)

haonan-li/cmmlu 공식 구현 pytorch

Tasks

Large Language Model

Similar Papers 제목 키워드 기반

IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

2025-01-27 · Sankalp KJ, Ashutosh Kumar, Laxmaan Balaji, Nikunj Kotecha 외

Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic…

BenchmarkingDiversityMMLUMulti-task Language Understanding

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

2024-02-20 · Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman 외

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Ar…

ArabicMMLULanguage Model EvaluationLanguage ModelingLanguage Modelling+2

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…

BnMMLU: Measuring Massive Multitask Language Understanding in Bengali

2025-05-25 · Saman Sarker Joy

The Massive Multitask Language Understanding (MMLU) benchmark has been widely used to evaluate language models across various domains. However, existing MMLU datasets primarily focus on high-resource languages such as En…

General KnowledgeMMLUMultiple-choice

Measuring Hong Kong Massive Multi-Task Language Understanding

2025-05-04 · Chuxue Cao, Zhenghao Zhu, Junqi Zhu, Guoying Lu 외

Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional …

MMLUMulti-task Language Understanding