paper-with-me

Papers

Measuring Massive Multitask Chinese Understanding

2023-04-25 · Hui Zeng

The development of large-scale Chinese language models is flourishing, yet there is a lack of corresponding capability assessments. Therefore, we propose a test to measure the multitask accuracy of large Chinese language models. This test encompasses four major domains, including medicine, law, psychology, and education, with 15 subtasks in medicine and 8 subtasks in education. We found that the best-performing models in the zero-shot setting outperformed the worst-performing models by nearly 18.6 percentage points on average. Across the four major domains, the highest average zero-shot accuracy of all models is 0.512. In the subdomains, only the GPT-3.5-turbo model achieved a zero-shot accuracy of 0.693 in clinical medicine, which was the highest accuracy among all models across all subtasks. All models performed poorly in the legal domain, with the highest zero-shot accuracy reaching only 0.239. By comprehensively evaluating the breadth and depth of knowledge across multiple disciplines, this test can more accurately identify the shortcomings of the models.

📄 PDF Abstract BibTeX arXiv:2304.12986

Code (2)

Felixgithub2017/MMCU 공식 구현
thudm/chatglm-6b pytorch

Tasks

All

Methods 이 논문이 사용한 방법론

{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Multi-Head Attention 설명 없음
Attention 설명 없음
Test 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음

Similar Papers 제목 키워드 기반

CMMLU: Measuring massive multitask language understanding in Chinese

2023-06-15 · Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang 외

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance becomes increasingly crucial and challenging. This paper aims to bridge this gap by introducing CMMLU, a comprehensive…

Large Language Model

KMMLU: Measuring Massive Multitask Language Understanding in Korean

2024-02-18 · Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim 외

We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmark…

kmmluLanguage Model EvaluationLanguage ModelingLanguage Modelling+1

BnMMLU: Measuring Massive Multitask Language Understanding in Bengali

2025-05-25 · Saman Sarker Joy

The Massive Multitask Language Understanding (MMLU) benchmark has been widely used to evaluate language models across various domains. However, existing MMLU datasets primarily focus on high-resource languages such as En…

General KnowledgeMMLUMultiple-choice

Measuring Massive Multitask Language Understanding

2020-09-07 · Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou 외

We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models mus…

Elementary MathematicsMulti-task Language UnderstandingMulti-Task LearningWorld Knowledge

An Improved Traditional Chinese Evaluation Suite for Foundation Model

2024-03-04 · Zhi-Rui Tam, Ya-Ting Pai, Yen-Wei Lee, Jun-Da Chen 외

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larg…

Multiple-choiceQuestion Answering