paper-with-me

홈 › Papers

CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data

2024-09-24 · Qian-Wen Zhang, Haochen Wang, Fang Li, Siyu An, Lingfeng Qiao, Liangcai Gao, Di Yin, Xing Sun

Online education platforms have significantly transformed the dissemination of educational resources by providing a dynamic and digital infrastructure. With the further enhancement of this transformation, the advent of Large Language Models (LLMs) has elevated the intelligence levels of these platforms. However, current academic benchmarks provide limited guidance for real-world industry scenarios. This limitation arises because educational applications require more than mere test question responses. To bridge this gap, we introduce CJEval, a benchmark based on Chinese Junior High School Exam Evaluations. CJEval consists of 26,136 samples across four application-level educational tasks covering ten subjects. These samples include not only questions and answers but also detailed annotations such as question types, difficulty levels, knowledge concepts, and answer explanations. By utilizing this benchmark, we assessed LLMs' potential applications and conducted a comprehensive analysis of their performance by fine-tuning on various educational tasks. Extensive experiments and discussions have highlighted the opportunities and challenges of applying LLMs in the field of education.

📄 PDF Abstract BibTeX arXiv:2409.16202

Code (1)

smilewhc/cjeval 공식 구현

Similar Papers 제목 키워드 기반

TCM-3CEval: A Triaxial Benchmark for Assessing Responses from Large Language Models in Traditional Chinese Medicine

2025-03-10 · Tianai Huang, Lu Lu, Jiayuan Chen, Lihao Liu 외

Large language models (LLMs) excel in various NLP tasks and modern medicine, but their evaluation in traditional Chinese medicine (TCM) is underexplored. To address this, we introduce TCM3CEval, a benchmark assessing LLM…

Decision Making

CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

2023-07-19 · Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu 외

With the rapid evolution of large language models (LLMs), there is a growing concern that they may pose risks or have negative social impacts. Therefore, evaluation of human values alignment is becoming increasingly impo…

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

2024-06-28 · Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin, Zhi-Long Ji 외

Due to the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Despite the datasets like MathVista proposed benchmarks for as…

DiversityMath

SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark

2023-07-27 · Liang Xu, Anqi Li, Lei Zhu, Hang Xue 외

Large language models (LLMs) have shown the potential to be integrated into human daily lives. Therefore, user preference is the most critical criterion for assessing LLMs' performance in real-world scenarios. However, e…

Language ModelingLanguage ModellingLarge Language Modelmodel

"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models

2025-02-17 · Jihao Gu, Yingyao Wang, Pi Bu, Chen Wang 외

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models' knowledge capacity and reliability. In this paper…

Object RecognitionQuestion AnsweringVisual Question Answering