paper-with-me

홈 › Papers

TCMBench: A Comprehensive Benchmark for Evaluating Large Language Models in Traditional Chinese Medicine

2024-06-03 · Wenjing Yue, Xiaoling Wang, Wei Zhu, Ming Guan, Huanran Zheng, Pengfei Wang, Changzhi Sun, Xin Ma

Large language models (LLMs) have performed remarkably well in various natural language processing tasks by benchmarking, including in the Western medical domain. However, the professional evaluation benchmarks for LLMs have yet to be covered in the traditional Chinese medicine(TCM) domain, which has a profound history and vast influence. To address this research gap, we introduce TCM-Bench, an comprehensive benchmark for evaluating LLM performance in TCM. It comprises the TCM-ED dataset, consisting of 5,473 questions sourced from the TCM Licensing Exam (TCMLE), including 1,300 questions with authoritative analysis. It covers the core components of TCMLE, including TCM basis and clinical practice. To evaluate LLMs beyond accuracy of question answering, we propose TCMScore, a metric tailored for evaluating the quality of answers generated by LLMs for TCM related questions. It comprehensively considers the consistency of TCM semantics and knowledge. After conducting comprehensive experimental analyses from diverse perspectives, we can obtain the following findings: (1) The unsatisfactory performance of LLMs on this benchmark underscores their significant room for improvement in TCM. (2) Introducing domain knowledge can enhance LLMs' performance. However, for in-domain models like ZhongJing-TCM, the quality of generated analysis text has decreased, and we hypothesize that their fine-tuning process affects the basic LLM capabilities. (3) Traditional metrics for text generation quality like Rouge and BertScore are susceptible to text length and surface semantic ambiguity, while domain-specific metrics such as TCMScore can further supplement and explain their evaluation results. These findings highlight the capabilities and limitations of LLMs in the TCM and aim to provide a more profound assistance to medical research.

📄 PDF Abstract BibTeX arXiv:2406.01126

Code (1)

ywjawmw/shennong-tcm-evaluation-benchmark

Tasks

BenchmarkingQuestion AnsweringText Generation

Similar Papers 제목 키워드 기반

TurkBench: A Benchmark for Evaluating Turkish Large Language Models

2026-01-11 · Çağrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz, Elif Ecem Arslan 외 arxiv

With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant progress has been made in evaluating Englis…

Instruction Following

VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model

2024-06-20 · Sibo Wang, Xiangkui Cao, Jie Zhang, Zheng Yuan 외

The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by concerns about biased outputs, a challenge…

Language ModelingLanguage Modelling

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

2026-04-22 · Yilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao 외 arxiv

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dim…

Machine Translation

Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

2023-11-27 · Munan Ning, Bin Zhu, Yujia Xie, Bin Lin 외

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goa…

Decision MakingQuestion Answering

ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models

2024-06-06 · Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang 외

Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. This development underscores the urgent need for evaluating value orientations and understanding of LLMs to …