paper-with-me

홈 › Papers

MedQA-CS: Benchmarking Large Language Models Clinical Skills Using an AI-SCE Framework

2024-10-02 · Zonghai Yao, Zihao Zhang, Chaolong Tang, Xingyu Bian, Youxia Zhao, Zhichao Yang, Junda Wang, Huixue Zhou, Won Seok Jang, Feiyun ouyang, Hong Yu

Artificial intelligence (AI) and large language models (LLMs) in healthcare require advanced clinical skills (CS), yet current benchmarks fail to evaluate these comprehensively. We introduce MedQA-CS, an AI-SCE framework inspired by medical education's Objective Structured Clinical Examinations (OSCEs), to address this gap. MedQA-CS evaluates LLMs through two instruction-following tasks, LLM-as-medical-student and LLM-as-CS-examiner, designed to reflect real clinical scenarios. Our contributions include developing MedQA-CS, a comprehensive evaluation framework with publicly available data and expert annotations, and providing the quantitative and qualitative assessment of LLMs as reliable judges in CS evaluation. Our experiments show that MedQA-CS is a more challenging benchmark for evaluating clinical skills than traditional multiple-choice QA benchmarks (e.g., MedQA). Combined with existing benchmarks, MedQA-CS enables a more comprehensive evaluation of LLMs' clinical capabilities for both open- and closed-source LLMs.

📄 PDF Abstract BibTeX arXiv:2410.01553

Code (1)

bio-nlp/medqa-cs 공식 구현

Tasks

BenchmarkingInstruction FollowingMedQAMultiple-choice

Similar Papers 제목 키워드 기반

What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs

2025-05-15 · Xinlan Yan, Di wu, Yibin Lei, Christof Monz 외

In this paper, we introduce S-MedQA, an English medical question-answering (QA) dataset for benchmarking large language models in fine-grained clinical specialties. We use S-MedQA to check the applicability of a popular …

AllBenchmarkingMedical Question AnsweringMedQA+1

HIVMedQA: Benchmarking large language models for HIV medical decision support

2025-07-24 · Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Roger Kouyos 외 arxiv

Large language models (LLMs) are emerging as valuable tools to support clinicians in routine decision-making. HIV management is a compelling use case due to its complexity, including diverse treatment options, comorbidit…

Prompt EngineeringQuestion Answering

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning

2025-07-21 · Sahana Srinivasan, Xuguang Ai, Thaddaeus Wai Soon Lo, Aidan Gilson 외 arxiv

Current benchmarks evaluating large language models (LLMs) in ophthalmology are limited in scope and disproportionately prioritise accuracy. We introduce BELO (BEnchmarking LLMs for Ophthalmology), a standardized and com…

Can large language models reason about medical questions?

2022-07-17 · Valentin Liévin, Christoffer Egeberg Hother, Andreas Geert Motzfeldt, Ole Winther

Although large language models (LLMs) often produce impressive outputs, it remains unclear how they perform in real-world scenarios requiring strong reasoning skills and expert domain knowledge. We set out to investigate…

MedQAMultiple-choiceMultiple Choice Question Answering (MCQA)Prompt Engineering+3

Large Language Models Encode Clinical Knowledge

2022-12-26 · Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi 외

Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but the quality bar for medical and clinical applications is high. Today, attempts to assess models…

Clinical KnowledgeMedQAMMLUMultiple-choice+4