paper-with-me

홈 › Papers

ZhuJiu: A Multi-dimensional, Multi-faceted Chinese Benchmark for Large Language Models

2023-08-28 · Baoli Zhang, Haining Xie, Pengfan Du, JunHao Chen, Pengfei Cao, Yubo Chen, Shengping Liu, Kang Liu, Jun Zhao

The unprecedented performance of large language models (LLMs) requires comprehensive and accurate evaluation. We argue that for LLMs evaluation, benchmarks need to be comprehensive and systematic. To this end, we propose the ZhuJiu benchmark, which has the following strengths: (1) Multi-dimensional ability coverage: We comprehensively evaluate LLMs across 7 ability dimensions covering 51 tasks. Especially, we also propose a new benchmark that focuses on knowledge ability of LLMs. (2) Multi-faceted evaluation methods collaboration: We use 3 different yet complementary evaluation methods to comprehensively evaluate LLMs, which can ensure the authority and accuracy of the evaluation results. (3) Comprehensive Chinese benchmark: ZhuJiu is the pioneering benchmark that fully assesses LLMs in Chinese, while also providing equally robust evaluation abilities in English. (4) Avoiding potential data leakage: To avoid data leakage, we construct evaluation data specifically for 37 tasks. We evaluate 10 current mainstream LLMs and conduct an in-depth discussion and analysis of their results. The ZhuJiu benchmark and open-participation leaderboard are publicly released at http://www.zhujiu-benchmark.com/ and we also provide a demo video at https://youtu.be/qypkJ89L1Ic.

📄 PDF Abstract BibTeX arXiv:2308.14353

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multifaceted Assessments of Traditional Chinese Word Segmentation Tool on Large Corpora

2022-11-01 · ROCLING 2022 11 · Wen-Chao Yeh, Yu-Lun Hsieh, Yung-Chun Chang, Wen-Lian Hsu

This study aims to evaluate three most popular word segmentation tool for a large Traditional Chinese corpus in terms of their efficiency, resource consumption, and cost. Specifically, we compare the performances of Jieb…

Chinese Word SegmentationGPUnamed-entity-recognitionNamed Entity Recognition+3

CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

2024-01-02 · Quan Tu, Shilong Fan, Zihang Tian, Rui Yan

Recently, the advent of large language models (LLMs) has revolutionized generative agents. Among them, Role-Playing Conversational Agents (RPCAs) attract considerable attention due to their ability to emotionally engage …

Tabular Two-Dimensional Correlation Analysis for Multifaceted Characterization Data

2023-11-27 · Shun Muroga, Satoshi Yamazaki, Koji Michishio, Hideaki Nakajima 외

We propose tabular two-dimensional correlation analysis for extracting features from multifaceted characterization data, essential for understanding material properties. This method visualizes similarities and phase lags…

Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation

2025-06-02 · Li Zhou, Lutong Yu, Dongchu Xie, Shaohuan Cheng 외

Culture is a rich and dynamic domain that evolves across both geography and time. However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often ove…

Multiple-choiceQuestion AnsweringVisual Question Answering

IJCNLP-2017 Task 2: Dimensional Sentiment Analysis for Chinese Phrases

2017-12-01 · IJCNLP 2017 12 · Liang-Chih Yu, Lung-Hao Lee, Jin Wang, Kam-Fai Wong

This paper presents the IJCNLP 2017 shared task on Dimensional Sentiment Analysis for Chinese Phrases (DSAP) which seeks to identify a real-value sentiment score of Chinese single words and multi-word phrases in the both…

Sentiment AnalysisTask 2