paper-with-me

홈 › Papers

CLUE: A Chinese Language Understanding Evaluation Benchmark

2020-04-13 · COLING 2020 8 · Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan

The advent of natural language understanding (NLU) benchmarks for English, such as GLUE and SuperGLUE allows new NLU models to be evaluated across a diverse set of tasks. These comprehensive benchmarks have facilitated a broad range of research and applications in natural language processing (NLP). The problem, however, is that most such benchmarks are limited to English, which has made it difficult to replicate many of the successes in English NLU for other languages. To help remedy this issue, we introduce the first large-scale Chinese Language Understanding Evaluation (CLUE) benchmark. CLUE is an open-ended, community-driven project that brings together 9 tasks spanning several well-established single-sentence/sentence-pair classification tasks, as well as machine reading comprehension, all on original Chinese text. To establish results on these tasks, we report scores using an exhaustive set of current state-of-the-art pre-trained Chinese models (9 in total). We also introduce a number of supplementary datasets and additional tools to help facilitate further progress on Chinese NLU. Our benchmark is released at https://www.CLUEbenchmarks.com

📄 PDF Abstract BibTeX arXiv:2004.05986

Code (3)

CLUEbenchmark/CLUE 공식 구현 tf
alibaba/EasyNLP jax
cbluebenchmark/cblue pytorch

Tasks

General ClassificationMachine Reading ComprehensionNatural Language UnderstandingReading ComprehensionSentenceSentence ClassificationSentence-Pair Classification

Similar Papers 제목 키워드 기반

Can Large Language Model Comprehend Ancient Chinese? A Preliminary Test on ACLUE

2023-10-14 · Yixuan Zhang, Haonan Li

Large language models (LLMs) have showcased remarkable capabilities in understanding and generating language. However, their ability in comprehending ancient languages, particularly ancient Chinese, remains largely unexp…

Language ModelingLanguage ModellingLarge Language Model

FewCLUE: A Chinese Few-shot Learning Evaluation Benchmark

2021-07-15 · Liang Xu, Xiaojing Lu, Chenyang Yuan, Xuanwei Zhang 외

Pretrained Language Models (PLMs) have achieved tremendous success in natural language understanding tasks. While different learning schemes -- fine-tuning, zero-shot, and few-shot learning -- have been widely explored a…

Few-Shot LearningMachine Reading ComprehensionNatural Language UnderstandingReading Comprehension+3

Shuo Wen Jie Zi: Rethinking Dictionaries and Glyphs for Chinese Language Pre-training

2023-05-30 · Yuxuan Wang, Jianghui Wang, Dongyan Zhao, Zilong Zheng

We introduce CDBERT, a new learning paradigm that enhances the semantics understanding ability of the Chinese PLMs with dictionary knowledge and structure of Chinese characters. We name the two core modules of CDBERT as …

Contrastive Learning

SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark

2023-07-27 · Liang Xu, Anqi Li, Lei Zhu, Hang Xue 외

Large language models (LLMs) have shown the potential to be integrated into human daily lives. Therefore, user preference is the most critical criterion for assessing LLMs' performance in real-world scenarios. However, e…

Language ModelingLanguage ModellingLarge Language Modelmodel

WYWEB: A NLP Evaluation Benchmark For Classical Chinese

2023-05-23 · Bo Zhou, Qianglong Chen, Tianyu Wang, Xiaomi Zhong 외

To fully evaluate the overall performance of different NLP models in a given domain, many evaluation benchmarks are proposed, such as GLUE, SuperGLUE and CLUE. The fi eld of natural language understanding has traditional…

Machine TranslationNatural Language UnderstandingReading ComprehensionSentence