paper-with-me

홈 › Papers

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

2026-06-17 · Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jiancheng Yang arxiv

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

📄 PDF Abstract BibTeX arXiv:2606.18613

Code (0)

등록된 구현이 없습니다.

Tasks

Clinical Knowledge

Similar Papers 제목 키워드 기반

Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care

2026-05-08 · Burcu Sayin, Ngoc Vo Hong, Ipek Baris Schlicht, Jacopo Staiano 외 arxiv

Clinical decision-making in emergency medicine demands rapid, accurate diagnoses under uncertainty. Despite benchmark progress, evidence for LLMs as interactive aids in live physician workflows remains sparse. MedSyn let…

Reverse Physician-AI Relationship: Full-process Clinical Diagnosis Driven by a Large Language Model

2025-08-14 · Shicheng Xu, Xin Huang, Zihao Wei, Liang Pang 외 arxiv

Full-process clinical diagnosis in the real world encompasses the entire diagnostic workflow that begins with only an ambiguous chief complaint. While artificial intelligence (AI), particularly large language models (LLM…

Improving Patient Pre-screening for Clinical Trials: Assisting Physicians with Large Language Models

2023-04-14 · Danny M. den Hamer, Perry Schoor, Tobias B. Polak, Daniel Kapitan

Physicians considering clinical trials for their patients are met with the laborious process of checking many text based eligibility criteria. Large Language Models (LLMs) have shown to perform well for clinical informat…

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

2026-05-07 · Maximillian Chen, Xuanming Zhang, Michael Peng, Zhou Yu 외 arxiv

The rise of Internet of Things (IoT) devices in the physical world necessitates voice-based interfaces capable of handling complex user experiences. While modern Large Language Models (LLMs) already demonstrate strong to…

Code Generation

Can LLMs Correct Physicians, Yet? Investigating Effective Interaction Methods in the Medical Domain

2024-03-29 · Burcu Sayin, Pasquale Minervini, Jacopo Staiano, Andrea Passerini

We explore the potential of Large Language Models (LLMs) to assist and potentially correct physicians in medical decision-making tasks. We evaluate several LLMs, including Meditron, Llama2, and Mistral, to analyze the ab…

Answer GenerationDecision Making