paper-with-me

홈 › Papers

CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios

2024-10-04 · Zetian Ouyang, Yishuai Qiu, LinLin Wang, Gerard de Melo, Ya zhang, Yanfeng Wang, Liang He

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench, a comprehensive benchmark with 14 expert-guided core clinical scenarios specifically designed to assess the medical ability of LLMs across 7 pivot dimensions. It comprises 33,735 questions derived from real-world medical reports of top-tier tertiary hospitals and authentic examination exercises. The reliability of this benchmark has been confirmed in several ways. Subsequent experiments with existing LLMs have led to the following findings: (i) Chinese medical LLMs underperform on this benchmark, especially where medical reasoning and factual consistency are vital, underscoring the need for advances in clinical knowledge and diagnostic accuracy. (ii) Several general-domain LLMs demonstrate substantial potential in medical clinics, while the limited input capacity of many medical LLMs hinders their practical use. These findings reveal both the strengths and limitations of LLMs in clinical scenarios and offer critical insights for medical research.

📄 PDF Abstract BibTeX arXiv:2410.03502

Code (1)

Optifine-TAT/CliMedBench 공식 구현

Tasks

Clinical KnowledgeDiagnostic

Similar Papers 제목 키워드 기반

LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language Models

2024-03-19 · Chuang Liu, Renren Jin, Yuqi Ren, Deyi Xiong

Chinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLM…

Multiple-choice

Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch

2025-02-24 · Xueru Wen, Jie Lou, Zichao Li, Yaojie Lu 외

Reward models (RMs) are crucial for aligning large language models (LLMs) with human preferences. However, most RM research is centered on English and relies heavily on synthetic resources, which leads to limited and les…

PromptCBLUE: A Chinese Prompt Tuning Benchmark for the Medical Domain

2023-10-22 · Wei Zhu, Xiaoling Wang, Huanran Zheng, Mosha Chen 외

Biomedical language understanding benchmarks are the driving forces for artificial intelligence applications with large language model (LLM) back-ends. However, most current benchmarks: (a) are limited to English which m…

Dialogue GenerationDialogue UnderstandingKnowledge ProbingLanguage Modeling+5

CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

2024-06-14 · Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Meijuan An 외

With the profound development of large language models(LLMs), their safety concerns have garnered increasing attention. However, there is a scarcity of Chinese safety benchmarks for LLMs, and the existing safety taxonomi…

Multiple-choiceQuestion Answering

CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs

2025-10-07 · Chengwei Wu, Jiapu Wang, Mingyang Gao, Xingrui Zhuo 외 arxiv

Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks. However, Chinese LLMs face unique challenges, primarily due to the dominance of unstructured free te…

Knowledge Graph CompletionQuestion AnsweringText Generation