paper-with-me

홈 › Papers

Reliable and diverse evaluation of LLM medical knowledge mastery

2024-09-22 · Yuxuan Zhou, Xien Liu, Chen Ning, Xiao Zhang, Ji Wu

Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical knowledge is still lacking. In the study, we propose a novel framework PretexEval that dynamically generates reliable and diverse test samples to evaluate LLMs for any given medical knowledge base. We notice that test samples produced directly from knowledge bases by templates or LLMs may introduce factual errors and also lack diversity. To address these issues, we introduce a novel schema into our proposed evaluation framework that employs predicate equivalence transformations to produce a series of variants for any given medical knowledge point. Finally, these produced predicate variants are converted into textual language, resulting in a series of reliable and diverse test samples to evaluate whether LLMs fully master the given medical factual knowledge point. Here, we use our proposed framework to systematically investigate the mastery of medical factual knowledge of 12 well-known LLMs, based on two knowledge bases that are crucial for clinical diagnosis and treatment. The evaluation results illustrate that current LLMs still exhibit significant deficiencies in fully mastering medical knowledge, despite achieving considerable success on some famous public benchmarks. These new findings provide valuable insights for developing medical-specific LLMs, highlighting that current LLMs urgently need to strengthen their comprehensive and in-depth mastery of medical knowledge before being applied to real-world medical scenarios.

📄 PDF Abstract BibTeX arXiv:2409.14302

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMedQA

Similar Papers 제목 키워드 기반

MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

2023-12-20 · Yan Cai, LinLin Wang, Ye Wang, Gerard de Melo 외

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive.…

Clinical KnowledgeDiagnostic

TCM-3CEval: A Triaxial Benchmark for Assessing Responses from Large Language Models in Traditional Chinese Medicine

2025-03-10 · Tianai Huang, Lu Lu, Jiayuan Chen, Lihao Liu 외

Large language models (LLMs) excel in various NLP tasks and modern medicine, but their evaluation in traditional Chinese medicine (TCM) is underexplored. To address this, we introduce TCM3CEval, a benchmark assessing LLM…

Decision Making

MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge

2024-06-05 · Yuxuan Zhou, Xien Liu, Chen Ning, Ji Wu

Large language models (LLMs) have excelled across domains, also delivering notable performance on the medical evaluation benchmarks, such as MedQA. However, there still exists a significant gap between the reported perfo…

MedQA

Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving

2025-06-10 · Yuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning 외

Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored. Inspired by Bloom's Taxonomy, we propose…

Counterfactual Monotonic Knowledge Tracing for Assessing Students' Dynamic Mastery of Knowledge Concepts

2023-08-07 · Moyu Zhang, Xinning Zhu, Chunhong Zhang, Wenchen Qian 외

As the core of the Knowledge Tracking (KT) task, assessing students' dynamic mastery of knowledge concepts is crucial for both offline teaching and online educational applications. Since students' mastery of knowledge co…

counterfactualKnowledge Tracing