paper-with-me

홈 › Papers

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?

2024-12-13 · Zhikai Lei, Tianyi Liang, Hanglei Hu, Jin Zhang, Yunhua Zhou, Yunfan Shao, Linyang Li, Chenchui Li, Changbo Wang, Hang Yan, Qipeng Guo

Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However, there is growing concern that LLMs may `game" these benchmarks due to data leakage, achieving high scores while struggling with tasks simple for humans. To substantively address the problem, we create GAOKAO-Eval, a comprehensive benchmark based on China's National College Entrance Examination (Gaokao), and conduct `closed-book" evaluations for representative models released prior to Gaokao. Contrary to prevailing consensus, even after addressing data leakage and comprehensiveness, GAOKAO-Eval reveals that high scores still fail to truly reflect human-aligned capabilities. To better understand this mismatch, We introduce the Rasch model from cognitive psychology to analyze LLM scoring patterns and identify two key discrepancies: 1) anomalous consistent performance across various question difficulties, and 2) high variance in performance on questions of similar difficulty. In addition, We identified inconsistent grading of LLM-generated answers among teachers and recurring mistake patterns. we find that the phenomenons are well-grounded in the motivations behind OpenAI o1, and o1's reasoning-as-difficulties can mitigate the mismatch. These results show that GAOKAO-Eval can reveal limitations in LLM capabilities not captured by current benchmarks and highlight the need for more LLM-aligned difficulty analysis.

📄 PDF Abstract BibTeX arXiv:2412.10056

Code (1)

open-compass/gaokao-eval 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

2023-05-21 · Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying 외

Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be…

Which is the Effective Way for Gaokao: Information Retrieval or Neural Networks?

2017-04-01 · EACL 2017 4 · Shangmin Guo, Xiangrong Zeng, Shizhu He, Kang Liu 외

As one of the most important test of China, Gaokao is designed to be difficult enough to distinguish the excellent high school students. In this work, we detailed the Gaokao History Multiple Choice Questions(GKHMC) and p…

Information RetrievalMultiple-choiceQuestion AnsweringReading Comprehension+2

Extract, Integrate, Compete: Towards Verification Style Reading Comprehension

2021-09-11 · Findings (EMNLP) 2021 11 · Chen Zhang, Yuxuan Lai, Yansong Feng, Dongyan Zhao

In this paper, we present a new verification style reading comprehension dataset named VGaokao from Chinese Language tests of Gaokao. Different from existing efforts, the new dataset is originally designed for native spe…

Reading Comprehension

GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation

2024-02-24 · Yi Zong, Xipeng Qiu

The Large Vision-Language Models (LVLMs) have demonstrated great abilities in image perception and language understanding. However, existing multimodal benchmarks focus on primary perception abilities and commonsense kno…

GCRC: A New Challenging MRC Dataset from Gaokao Chinese for Explainable Evaluation

2021-08-01 · Findings (ACL) 2021 8 · Hongye Tan, Xiaoyue Wang, Yu Ji, Ru Li 외