paper-with-me

Papers

ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis

2025-02-16 · Xu Wang, Jiaju Kang, Puyu Han, Yubao Zhao, Qian Liu, Liwenfei He, Lingqiong Zhang, Lingyun Dai, Yongcheng Wang, Jie Tao

We present ECG-Expert-QA, a comprehensive multimodal dataset for evaluating diagnostic capabilities in electrocardiogram (ECG) interpretation. It combines real-world clinical ECG data with systematically generated synthetic cases, covering 12 essential diagnostic tasks and totaling 47,211 expert-validated QA pairs. These encompass diverse clinical scenarios, from basic rhythm recognition to complex diagnoses involving rare conditions and temporal changes. A key innovation is the support for multi-turn dialogues, enabling the development of conversational medical AI systems that emulate clinician-patient or interprofessional interactions. This allows for more realistic assessment of AI models' clinical reasoning, diagnostic accuracy, and knowledge integration. Constructed through a knowledge-guided framework with strict quality control, ECG-Expert-QA ensures linguistic and clinical consistency, making it a high-quality resource for advancing AI-assisted ECG interpretation. It challenges models with tasks like identifying subtle ischemic changes and interpreting complex arrhythmias in context-rich scenarios. To promote research transparency and collaboration, the dataset, accompanying code, and prompts are publicly released at https://github.com/Zaozzz/ECG-Expert-QA

📄 PDF Abstract BibTeX arXiv:2502.17475

Code (1)

zaozzz/ecg-expert-qa 공식 구현

Tasks

DiagnosticRhythm

Similar Papers 제목 키워드 기반

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

2026-04-07 · Han Jang, Junhyeok Lee, Heeseong Eum, Kyu Sung Choi arxiv

Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are predominantly trained on professional literature, limiting their ability to…

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

2025-07-10 · Shuang Zhou, Wenya Xie, Jiaxi Li, Zaifu Zhan 외 arxiv

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasonin…

MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

2026-06-19 · Han Jang, Junhyeok Lee, Songsoo Kim, Chae Young Lim 외 arxiv

Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains underexplored. With the 21st Century Cures Act now mandating immedi…

LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation

2025-06-04 · Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha 외

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehe…

Multiple-choice

A Benchmark for Long-Form Medical Question Answering

2024-11-14 · Pedram Hosseini, Jessica M. Sin, Bing Ren, Bryceton G. Thomas 외

There is a lack of benchmarks for evaluating large language models (LLMs) in long-form medical question answering (QA). Most existing medical QA evaluation benchmarks focus on automatic metrics and multiple-choice questi…

Answer GenerationFormMedical Question AnsweringMultiple-choice+1