paper-with-me

홈 › Papers

EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning

2026-06-22 · Yitong Qiao, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu, Kui Ren arxiv

Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often operating on idealized, clean EHRs via static SQL generation rather than interactive execution. In this work, we introduce EHR-Complex, a large-scale benchmark designed for interactive clinical database reasoning. Built on the large MIMIC-IV substrate (365K patients, 31 tables, 500M+ records), EHR-Complex comprises about 52K tasks spanning six clinical intents, supporting both patient-level and population-level queries, where each task requires an agent to interact with a sandboxed environment by executing SQL queries or Python code. Notably, EHR-Complex considers the real-world SQL task complexity for longitudinal multi-table aggregation and compositional reasoning, resulting in 31.93 SQL structural components per query on average. Evaluation results on EHR-Complex reveal the clinical difficulty of these EHR reasoning scenarios, with the top-performing model achieving only 62.3% exact-match accuracy. Pass^k consistency drops below 50% for nearly all evaluated models at k=4, exposing broad stochastic fragility. A fine-grained analysis of more than 3,800 failed trajectories for representative LLMs reveals three dominant failure modes: SQL logic errors, medical-code lookup failures, and semantic misunderstandings. EHR-Complex provides a rigorous testbed for clinical agents and highlights remaining gaps in robust reasoning for large-scale EHR analysis.

📄 PDF Abstract BibTeX arXiv:2606.23301

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare

2026-05-12 · Yihao Wang, Haoran Xu, Renjie Gu, Yixuan Ye 외 arxiv

The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, existing benchmarks primarily focus on dai…

Medchain: Bridging the Gap Between LLM Agents and Clinical Practice through Interactive Sequential Benchmarking

2024-12-02 · Jie Liu, Wenxuan Wang, Zizhan Ma, Guolin Huang 외

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have …

BenchmarkingDecision MakingLanguage ModelingLanguage Modelling+3

MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning

2025-03-10 · Xiangru Tang, Daniel Shao, Jiwoong Sohn, Jiapeng Chen 외

Large Language Models (LLMs) have shown impressive performance on existing medical question-answering benchmarks. This high performance makes it increasingly difficult to meaningfully evaluate and differentiate advanced …

BenchmarkingMedical Question AnsweringQuestion Answering

Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

2024-02-28 · Hanjie Chen, Zhouxiang Fang, Yash Singla, Mark Dredze

LLMs have demonstrated impressive performance in answering medical questions, such as achieving passing scores on medical licensing examinations. However, medical board exams or general clinical questions do not capture …

BenchmarkingMultiple-choiceQuestion Answering

Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios

2024-11-16 · Shaochen Xu, Yifan Zhou, Zhengliang Liu, Zihao Wu 외

Artificial Intelligence (AI) has become essential in modern healthcare, with large language models (LLMs) offering promising advances in clinical decision-making. Traditional model-based approaches, including those lever…

Action GenerationAI AgentDecision MakingDiagnostic+2