paper-with-me

홈 › Papers

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

2025-05-16 · Xiaomin Li, Mingye Gao, Yuexing Hao, Taoran Li, Guangya Wan, Zihan Wang, Yijun Wang

Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making. However, it remains unclear whether Large Language Models (LLMs) can reliably follow such structured protocols. In this work, we introduce MedGUIDE, a new benchmark for evaluating LLMs on their ability to make guideline-consistent clinical decisions. MedGUIDE is constructed from 55 curated NCCN decision trees across 17 cancer types and uses clinical scenarios generated by LLMs to create a large pool of multiple-choice diagnostic questions. We apply a two-stage quality selection process, combining expert-labeled reward models and LLM-as-a-judge ensembles across ten clinical and linguistic criteria, to select 7,747 high-quality samples. We evaluate 25 LLMs spanning general-purpose, open-source, and medically specialized models, and find that even domain-specific LLMs often underperform on tasks requiring structured guideline adherence. We also test whether performance can be improved via in-context guideline inclusion or continued pretraining. Our findings underscore the importance of MedGUIDE in assessing whether LLMs can operate safely within the procedural frameworks expected in real-world clinical settings.

📄 PDF Abstract BibTeX arXiv:2505.11613

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDecision MakingDiagnosticMultiple-choice

Similar Papers 제목 키워드 기반

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

2026-05-26 · Yuhao Shen, Lang Cao, Simo Du, Yuqing Wang 외 arxiv

Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, and recommendation rules. However, existing methods often use CPGs as …

VL-MedGuide: A Visual-Linguistic Large Model for Intelligent and Explainable Skin Disease Auxiliary Diagnosis

2025-08-08 · Kexin Yu, Zihan Xu, Jialei Xie, Carter Adams arxiv

Accurate diagnosis of skin diseases remains a significant challenge due to the complex and diverse visual features present in dermatoscopic images, often compounded by a lack of interpretability in existing purely visual…

Prompt Engineering

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

2026-07-10 · Yanzhen Chen, Zihan Xu, Xiaocheng Zhang, Zhiting Fan 외 arxiv

In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and to…

Information Retrieval

GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space

2026-08-31 · Boqi Chen, Xudong Liu, Yunke Ao, Heejin Do 외 arxiv

Large Language Models (LLMs) show great potential as clinical agents, yet existing benchmarks reduce clinical workflows to static predictions or unconstrained Markov Decision Processes (MDPs) with coarse action sets. To …

MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology

2025-11-25 · Kiril Vasilev, Alexandre Misrahi, Eeshaan Jain, Phil F Cheng 외 arxiv

Multimodal Large Language Models (LLMs) hold promise for biomedical reasoning, but current benchmarks fail to capture the complexity of real-world clinical workflows. Existing evaluations primarily assess unimodal, decon…