paper-with-me

Papers

MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge

2024-06-05 · Yuxuan Zhou, Xien Liu, Chen Ning, Ji Wu

Large language models (LLMs) have excelled across domains, also delivering notable performance on the medical evaluation benchmarks, such as MedQA. However, there still exists a significant gap between the reported performance and the practical effectiveness in real-world medical scenarios. In this paper, we aim to explore the causes of this gap by employing a multifaceted examination schema to systematically probe the actual mastery of medical knowledge by current LLMs. Specifically, we develop a novel evaluation framework MultifacetEval to examine the degree and coverage of LLMs in encoding and mastering medical knowledge at multiple facets (comparison, rectification, discrimination, and verification) concurrently. Based on the MultifacetEval framework, we construct two multifaceted evaluation datasets: MultiDiseK (by producing questions from a clinical disease knowledge base) and MultiMedQA (by rephrasing each question from a medical benchmark MedQA into multifaceted questions). The experimental results on these multifaceted datasets demonstrate that the extent of current LLMs in mastering medical knowledge is far below their performance on existing medical benchmarks, suggesting that they lack depth, precision, and comprehensiveness in mastering medical knowledge. Consequently, current LLMs are not yet ready for application in real-world medical tasks. The codes and datasets are available at https://github.com/THUMLP/MultifacetEval.

📄 PDF Abstract BibTeX arXiv:2406.02919

Code (1)

thumlp/multifaceteval 공식 구현 pytorch

Tasks

MedQA

Similar Papers 제목 키워드 기반

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

2025-05-22 · Kai Li, Can Shen, Yile Liu, Jirui Han 외

The rapid advancement and expanding applications of Audio Large Language Models (ALLMs) demand a rigorous understanding of their trustworthiness. However, systematic research on evaluating these models, particularly conc…

BenchmarkingFairnessHallucination

ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models

2024-06-06 · Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang 외

Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. This development underscores the urgent need for evaluating value orientations and understanding of LLMs to …

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark

2026-06-01 · Seonghyeon Go, Yumin Kim arxiv

As generative platforms such as Suno and Udio reach human-grade audio quality, the scope of AI's utility has expanded across the entire music production workflow. Beyond simple track generation, these advancements have c…

Binary Classification

Can Large Language Models Master Complex Card Games?

2025-09-01 · Wei Wang, Fuqing Bie, Junzhe Chen, Dan Zhang 외 arxiv

Complex games have long been an important benchmark for testing the progress of artificial intelligence algorithms. AlphaGo, AlphaZero, and MuZero have defeated top human players in Go and Chess, garnering widespread soc…

Dynamic Evaluation of Large Language Models by Meta Probing Agents

2024-02-21 · Kaijie Zhu, Jindong Wang, Qinlin Zhao, Ruochen Xu 외

Evaluation of large language models (LLMs) has raised great concerns in the community due to the issue of data contamination. Existing work designed evaluation protocols using well-defined algorithms for specific tasks, …

Data Augmentation