paper-with-me

홈 › Papers

GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation

2024-02-24 · Yi Zong, Xipeng Qiu

The Large Vision-Language Models (LVLMs) have demonstrated great abilities in image perception and language understanding. However, existing multimodal benchmarks focus on primary perception abilities and commonsense knowledge which are insufficient to reflect the comprehensive capabilities of LVLMs. We propose GAOKAO-MM, a multimodal benchmark based on the Chinese College Entrance Examination (GAOKAO), comprising of 8 subjects and 12 types of images, such as diagrams, function graphs, maps and photos. GAOKAO-MM derives from native Chinese context and sets human-level requirements for the model's abilities, including perception, understanding, knowledge and reasoning. We evaluate 10 LVLMs and find that the accuracies of all of them are lower than 50%, with GPT-4-Vison (48.1%), Qwen-VL-Plus (41.2%) and Gemini-Pro-Vision (35.1%) ranking in the top three positions. The results of our multi-dimension analysis indicate that LVLMs have moderate distance towards Artificial General Intelligence (AGI) and provide insights facilitating the development of multilingual LVLMs.

📄 PDF Abstract BibTeX arXiv:2402.15745

Code (1)

openmoss/gaokao-mm 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

2023-05-21 · Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying 외

Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be…

GCRC: A New Challenging MRC Dataset from Gaokao Chinese for Explainable Evaluation

2021-08-01 · Findings (ACL) 2021 8 · Hongye Tan, Xiaoyue Wang, Yu Ji, Ru Li 외

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?

2024-12-13 · Zhikai Lei, Tianyi Liang, Hanglei Hu, Jin Zhang 외

Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However, there is growing concern that LLMs may…

Extract, Integrate, Compete: Towards Verification Style Reading Comprehension

2021-09-11 · Findings (EMNLP) 2021 11 · Chen Zhang, Yuxuan Lai, Yansong Feng, Dongyan Zhao

In this paper, we present a new verification style reading comprehension dataset named VGaokao from Chinese Language tests of Gaokao. Different from existing efforts, the new dataset is originally designed for native spe…

Reading Comprehension

One-shot Learning for Question-Answering in Gaokao History Challenge

2018-06-24 · COLING 2018 8 · Zhuosheng Zhang, Hai Zhao

Answering questions from university admission exams (Gaokao in Chinese) is a challenging AI task since it requires effective representation to capture complicated semantic relations between questions and answers. In this…

One-Shot LearningQuestion Answering