paper-with-me

홈 › Papers

GCRC: A New Challenging MRC Dataset from Gaokao Chinese for Explainable Evaluation

2021-08-01 · Findings (ACL) 2021 8 · Hongye Tan, Xiaoyue Wang, Yu Ji, Ru Li, XiaoLi Li, Zhiwei Hu, Yunxiao Zhao, Xiaoqi Han
📄 PDF Abstract BibTeX

Code (1)

sxunlp/gcrc 공식 구현

Similar Papers 제목 키워드 기반

GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation

2024-02-24 · Yi Zong, Xipeng Qiu

The Large Vision-Language Models (LVLMs) have demonstrated great abilities in image perception and language understanding. However, existing multimodal benchmarks focus on primary perception abilities and commonsense kno…

Extract, Integrate, Compete: Towards Verification Style Reading Comprehension

2021-09-11 · Findings (EMNLP) 2021 11 · Chen Zhang, Yuxuan Lai, Yansong Feng, Dongyan Zhao

In this paper, we present a new verification style reading comprehension dataset named VGaokao from Chinese Language tests of Gaokao. Different from existing efforts, the new dataset is originally designed for native spe…

Reading Comprehension

One-shot Learning for Question-Answering in Gaokao History Challenge

2018-06-24 · COLING 2018 8 · Zhuosheng Zhang, Hai Zhao

Answering questions from university admission exams (Gaokao in Chinese) is a challenging AI task since it requires effective representation to capture complicated semantic relations between questions and answers. In this…

One-Shot LearningQuestion Answering

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

2023-05-21 · Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying 외

Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be…

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?

2024-12-13 · Zhikai Lei, Tianyi Liang, Hanglei Hu, Jin Zhang 외

Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However, there is growing concern that LLMs may…