paper-with-me

Papers

The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability

2025-09-29 · Linlu Gong, Ante Wang, Yunghwei Lai, Weizhi Ma, Yang Liu arxiv

An effective physician should possess a combination of empathy, expertise, patience, and clear communication when treating a patient. Recent advances have successfully endowed AI doctors with expert diagnostic skills, particularly the ability to actively seek information through inquiry. However, other essential qualities of a good doctor remain overlooked. To bridge this gap, we present MAQuE(Medical Agent Questioning Evaluation), the largest-ever benchmark for the automatic and comprehensive evaluation of medical multi-turn questioning. It features 3,000 realistically simulated patient agents that exhibit diverse linguistic patterns, cognitive limitations, emotional responses, and tendencies for passive disclosure. We also introduce a multi-faceted evaluation framework, covering task success, inquiry proficiency, dialogue competence, inquiry efficiency, and patient experience. Experiments on different LLMs reveal substantial challenges across the evaluation aspects. Even state-of-the-art models show significant room for improvement in their inquiry capabilities. These models are highly sensitive to variations in realistic patient behavior, which considerably impacts diagnostic accuracy. Furthermore, our fine-grained metrics expose trade-offs between different evaluation perspectives, highlighting the challenge of balancing performance and practicality in real-world clinical settings.

📄 PDF Abstract BibTeX arXiv:2509.24958

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive Annotations

2024-10-18 · Vishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu 외

Medical task-oriented dialogue systems can assist doctors by collecting patient medical history, aiding in diagnosis, or guiding treatment selection, thereby reducing doctor burnout and expanding access to medical servic…

Natural Language UnderstandingTask-Oriented Dialogue SystemsText Generation

Doctor XAvIer: Explainable Diagnosis on Physician-Patient Dialogues and XAI Evaluation

2022-04-11 · BioNLP (ACL) 2022 5 · Hillary Ngai, Frank Rudzicz

We introduce Doctor XAvIer, a BERT-based diagnostic system that extracts relevant clinical data from transcribed patient-doctor dialogues and explains predictions using feature attribution methods. We present a novel per…

ClassificationDiagnosticExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)+6

Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions

2025-03-28 · Mohammad Almansoori, Komal Kumar, Hisham Cholakkal

In this work, we introduce MedAgentSim, an open-source simulated clinical environment with doctor, patient, and measurement agents designed to evaluate and enhance LLM performance in dynamic diagnostic settings. Unlike p…

Diagnostic

Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents

2023-10-13 · Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon 외

Human-like chatbots necessitate the use of commonsense reasoning in order to effectively comprehend and respond to implicit information present within conversations. Achieving such coherence and informativeness in respon…

InformativenessKnowledge DistillationResponse Generation

Two eyes, Two views, and finally, One summary! Towards Multi-modal Multi-tasking Knowledge-Infused Medical Dialogue Summarization

2024-07-21 · Anisha Saha, Abhisek Tiwari, Sai Ruthvik, Sriparna Saha

We often summarize a multi-party conversation in two stages: chunking with homogeneous units and summarizing the chunks. Thus, we hypothesize that there exists a correlation between homogeneous speaker chunking and overa…

ChunkingConversation Summarizationdialogue summary