GCRC: A New Challenging MRC Dataset from Gaokao Chinese for Explainable Evaluation
Code (1)
Similar Papers 제목 키워드 기반
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
The Large Vision-Language Models (LVLMs) have demonstrated great abilities in image perception and language understanding. However, existing multimodal benchmarks focus on primary perception abilities and commonsense kno…
Extract, Integrate, Compete: Towards Verification Style Reading Comprehension
In this paper, we present a new verification style reading comprehension dataset named VGaokao from Chinese Language tests of Gaokao. Different from existing efforts, the new dataset is originally designed for native spe…
Reading ComprehensionOne-shot Learning for Question-Answering in Gaokao History Challenge
Answering questions from university admission exams (Gaokao in Chinese) is a challenging AI task since it requires effective representation to capture complicated semantic relations between questions and answers. In this…
One-Shot LearningQuestion AnsweringEvaluating the Performance of Large Language Models on GAOKAO Benchmark
Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be…
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However, there is growing concern that LLMs may…