paper-with-me

홈 › Papers

CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark

2021-12-27 · Yuan YAO, Qingxiu Dong, Jian Guan, Boxi Cao, Zhengyan Zhang, Chaojun Xiao, Xiaozhi Wang, Fanchao Qi, Junwei Bao, Jinran Nie, Zheni Zeng, Yuxian Gu, Kun Zhou, Xuancheng Huang, Wenhao Li, Shuhuai Ren, Jinliang Lu, Chengqiang Xu, Huadong Wang, Guoyang Zeng, Zile Zhou, Jiajun Zhang, Juanzi Li, Minlie Huang, Rui Yan, Xiaodong He, Xiaojun Wan, Xin Zhao, Xu sun, Yang Liu, Zhiyuan Liu, Xianpei Han, Erhong Yang, Zhifang Sui, Maosong Sun

Realizing general-purpose language intelligence has been a longstanding goal for natural language processing, where standard evaluation benchmarks play a fundamental and guiding role. We argue that for general-purpose language intelligence evaluation, the benchmark itself needs to be comprehensive and systematic. To this end, we propose CUGE, a Chinese Language Understanding and Generation Evaluation benchmark with the following features: (1) Hierarchical benchmark framework, where datasets are principally selected and organized with a language capability-task-dataset hierarchy. (2) Multi-level scoring strategy, where different levels of model performance are provided based on the hierarchical framework. To facilitate CUGE, we provide a public leaderboard that can be customized to support flexible model judging criteria. Evaluation results on representative pre-trained language models indicate ample room for improvement towards general-purpose language intelligence. CUGE is publicly available at cuge.baai.ac.cn.

📄 PDF Abstract BibTeX arXiv:2112.13610

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and Benchmark

2023-02-18 · Dakuan Lu, Hengkui Wu, Jiaqing Liang, Yipei Xu 외

To advance Chinese financial natural language processing (NLP), we introduce BBT-FinT5, a new Chinese financial pre-training language model based on the T5 model. To support this effort, we have built BBT-FinCorpus, a la…

Language ModelingLanguage Modelling

YACLC: A Chinese Learner Corpus with Multidimensional Annotation

2021-12-30 · Yingying Wang, Cunliang Kong, Liner Yang, Yijun Wang 외

Learner corpus collects language data produced by L2 learners, that is second or foreign-language learners. This resource is of great relevance for second language acquisition research, foreign-language teaching, and aut…

Grammatical Error CorrectionLanguage AcquisitionSentence

Can Pre-trained Language Models Understand Chinese Humor?

2024-07-04 · Yuyan Chen, Zhixu Li, Jiaqing Liang, Yanghua Xiao 외

Humor understanding is an important and challenging research in natural language processing. As the popularity of pre-trained language models (PLMs), some recent work makes preliminary attempts to adopt PLMs for humor re…

cuGenOpt: A GPU-Accelerated General-Purpose Metaheuristic Framework for Combinatorial Optimization

2026-03-19 · Yuyang Liu arxiv

Combinatorial optimization problems arise in logistics, scheduling, and resource allocation, yet existing approaches face a fundamental trade-off among generality, performance, and usability. We present cuGenOpt, a GPU-a…

Fùxì: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation

2025-03-20 · Shangqing Zhao, Yuhao Zhou, Yupei Ren, Zhe Chen 외

Ancient Chinese text processing presents unique challenges for large language models (LLMs) due to its distinct linguistic features, complex structural constraints, and rich cultural context. While existing benchmarks ha…

Multiple-choiceText Generation