paper-with-me

홈 › Papers

Large Language Models in the Clinic: A Comprehensive Benchmark

2024-04-25 · Fenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng, Haoming Jiang, Yifan Gao, Priyanka Nigam, Sreyashi Nag, Bing Yin, Yining Hua, Xuan Zhou, Omid Rohanian, Anshul Thakur, Lei Clifton, David A. Clifton

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the close-ended question-answering (QA) task with answer options for evaluation. However, many clinical decisions involve answering open-ended questions without pre-set options. To better understand LLMs in the clinic, we construct a benchmark ClinicBench. We first collect eleven existing datasets covering diverse clinical language generation, understanding, and reasoning tasks. Furthermore, we construct six novel datasets and clinical tasks that are complex but common in real-world practice, e.g., open-ended decision-making, long document processing, and emerging drug analysis. We conduct an extensive evaluation of twenty-two LLMs under both zero-shot and few-shot settings. Finally, we invite medical experts to evaluate the clinical usefulness of LLMs. The benchmark data is available at https://github.com/AI-in-Health/ClinicBench.

📄 PDF Abstract BibTeX arXiv:2405.00716

Code (1)

ai-in-health/clinicbench 공식 구현

Tasks

Decision MakingDocument SummarizationQuestion AnsweringText Generation

Similar Papers 제목 키워드 기반

MedQA-CS: Benchmarking Large Language Models Clinical Skills Using an AI-SCE Framework

2024-10-02 · Zonghai Yao, Zihao Zhang, Chaolong Tang, Xingyu Bian 외

Artificial intelligence (AI) and large language models (LLMs) in healthcare require advanced clinical skills (CS), yet current benchmarks fail to evaluate these comprehensively. We introduce MedQA-CS, an AI-SCE framework…

BenchmarkingInstruction FollowingMedQAMultiple-choice

CLIMB: A Benchmark of Clinical Bias in Large Language Models

2024-07-07 · Yubo Zhang, Shudi Hou, Mingyu Derek Ma, Wei Wang 외

Large language models (LLMs) are increasingly applied to clinical decision-making. However, their potential to exhibit bias poses significant risks to clinical equity. Currently, there is a lack of benchmarks that system…

counterfactualDecision Making

CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios

2024-10-04 · Zetian Ouyang, Yishuai Qiu, LinLin Wang, Gerard de Melo 외

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We p…

Clinical KnowledgeDiagnostic

Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization

2025-11-29 · Boyang Gu, Hongjian Zhou, Bradley Max Segal, Jinge Wu 외 arxiv

Recent advances in large language models (LLMs) have shown strong reasoning capabilities through large-scale pretraining and post-training reinforcement learning, demonstrated by DeepSeek-R1. However, current post-traini…

Reinforcement Learning

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

2026-04-16 · Meng-Xun Li, Wen-Hui Deng, Zhi-Xing Wu, Chun-Xiao Jin 외 arxiv

Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains largely underexplored due to the lack of fine-grained, annotated dat…

Multi-Label ClassificationVisual Question AnsweringImage Captioning