paper-with-me

홈 › Papers

MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction

2026-04-05 · Zhichao Yang, Gregory D. Lyng, Sanjit Singh Batra, Robert E. Tillman arxiv

Medical concept extraction from electronic health records underpins many downstream applications, yet remains challenging because medically meaningful concepts are frequently implied rather than explicitly stated in medical narratives. Existing benchmarks with human-annotated evidence spans underscore the importance of grounding extracted concepts in medical text. However, they predominantly focus on explicitly stated concepts instead of implicit concepts. We present MedicalBench, a benchmark for medical concept extraction with evidence grounding that evaluates implicit medical reasoning. MedicalBench formulates medical concept extraction as a verification task over medical note-concept pairs, coupled with sentence-level evidence identification. Built from MIMIC-IV discharge summaries and human-verified ICD-10 codes, the dataset is curated through a multi-stage large language model (LLM) triage pipeline followed by medical annotation and expert review. It deliberately includes implicit positives, semantically confusable negatives, and cases where LLM judgments disagree with medical expert assessments. We define two complementary evaluation tasks: (1) medical concept extraction and (2) sentence-level evidence retrieval, enabling assessment of both correctness and interpretability. Benchmarking state-of-the-art LLMs reveals that performance remains modest, highlighting the difficulty of extracting implicitly expressed concepts. We further show that performance is largely invariant to note length, indicating that MedicalBench isolates reasoning difficulty rather than superficial confounders. MedicalBench provides the first systematic benchmark for implicit, evidence-grounded medical concept extraction, offering a foundation for developing medical language models that can both identify medically relevant concepts and justify their predictions in a transparent and medically faithful manner.

📄 PDF Abstract BibTeX arXiv:2605.20197

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

2022-06-30 · Jason Alan Fries, Leon Weber, Natasha Seelam, Gabriel Altay 외

Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-s…

DiversityLanguage Model EvaluationLanguage ModelingLanguage Modelling+4

Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam

2025-07-26 · Cesar Augusto Madid Truyts, Amanda Gomes Rabelo, Gabriel Mesquita de Souza, Daniel Scaldaferri Lages 외 arxiv

Artificial intelligence (AI) has shown the potential to revolutionize healthcare by improving diagnostic accuracy, optimizing workflows, and personalizing treatment plans. Large Language Models (LLMs) and Multimodal Larg…

Multimodal Reasoning

Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics

2023-04-01 · Jason Holmes, Zhengliang Liu, Lian Zhang, Yuzhen Ding 외

We present the first study to investigate Large Language Models (LLMs) in answering radiation oncology physics questions. Because popular exams like AP Physics, LSAT, and GRE have large test-taker populations and ample t…

Evaluating large language models in medical applications: a survey

2024-05-13 · Xiaolan Chen, Jiayang Xiang, Shanfu Lu, Yexin Liu 외

Large language models (LLMs) have emerged as powerful tools with transformative potential across numerous domains, including healthcare and medicine. In the medical domain, LLMs hold promise for tasks ranging from clinic…

Survey

Benchmarking Large Language Models on CMExam - A comprehensive Chinese Medical Exam Dataset

2023-09-26 · NeurIPS 2023 11

Recent advancements in large language models (LLMs) have transformed the field of question answering (QA). However, evaluating LLMs in the medical field is challenging due to the lack of standardized and comprehensive da…