paper-with-me

Papers

An Improved Traditional Chinese Evaluation Suite for Foundation Model

2024-03-04 · Zhi-Rui Tam, Ya-Ting Pai, Yen-Wei Lee, Jun-Da Chen, Wei-Min Chu, Sega Cheng, Hong-Han Shuai

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larger and boasts a more balanced subject distribution than its predecessor, Taiwan Massive Multitask Language Understanding (TMMLU). We also benchmark closed-source models and 26 open-weight Chinese large language models (LLMs) of parameters ranging from 1.8B to 72B on the proposed TMMLU+. Our findings reveal that (1.) Traditional Chinese models still trail behind their Simplified Chinese counterparts, highlighting a need for more focused advancements in LLMs catering to Traditional Chinese. (2.) Current LLMs still fall short of human performance in average scores, indicating a potential need for future research to delve deeper into social science and humanities subjects. (3.) Among all the tokenization compression metrics examined, we identify that only the fertility score uniquely demonstrates strong correlations with our benchmark results. We foresee that TMMLU+ will pinpoint areas for future model improvement, thereby narrowing the gap between machine and human linguistic capabilities and supporting researchers in developing Traditional Chinese LLMs. Our dataset, along with the benchmark source code, is accessible at huggingface.co/datasets/ikala/tmmluplus.

📄 PDF Abstract BibTeX arXiv:2403.01858

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceQuestion Answering

Similar Papers 제목 키워드 기반

C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

2023-05-15 · NeurIPS 2023 11 · Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang 외

New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-Eval, the first comprehensive Chinese evaluation suite designed to assess advanced knowledge and re…

Multiple-choice

VisTai: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan

2025-03-13 · Zhi Rui Tam, Ya-Ting Pai, Yen-Wei Lee

In this paper, we propose a comprehensive evaluation benchmark for Visual Language Models (VLM) in Traditional Chinese. Our evaluation suite, the first of its kind, contains two complementary components: (1) VisTai-MCQ, …

BenchmarkingDialogue Generation

Advancing the Evaluation of Traditional Chinese Language Models: Towards a Comprehensive Benchmark Suite

2023-09-15 · Chan-Jan Hsu, Chang-Le Liu, Feng-Ting Liao, Po-chun Hsu 외

The evaluation of large language models is an essential task in the field of language understanding and generation. As language models continue to advance, the need for effective benchmarks to assess their performance ha…

Question Answering

Rethinking the Foundations for Continual Reinforcement Learning

2025-04-10 · Michael Bowling, Esraa Elelimy

Algorithms and approaches for continual reinforcement learning have gained increasing attention. Much of this early progress rests on the foundations and standard practices of traditional reinforcement learning, without …

Continual Learningreinforcement-learningReinforcement Learning

Eval-GCSC: A New Metric for Evaluating ChatGPT's Performance in Chinese Spelling Correction

2023-11-14 · Kunting Li, Yong Hu, Shaolei Wang, Hanhan Ma 외

ChatGPT has demonstrated impressive performance in various downstream tasks. However, in the Chinese Spelling Correction (CSC) task, we observe a discrepancy: while ChatGPT performs well under human evaluation, it scores…

Semantic SimilaritySemantic Textual SimilaritySpelling Correction