paper-with-me

홈 › Papers

Can LLMs Correct Physicians, Yet? Investigating Effective Interaction Methods in the Medical Domain

2024-03-29 · Burcu Sayin, Pasquale Minervini, Jacopo Staiano, Andrea Passerini

We explore the potential of Large Language Models (LLMs) to assist and potentially correct physicians in medical decision-making tasks. We evaluate several LLMs, including Meditron, Llama2, and Mistral, to analyze the ability of these models to interact effectively with physicians across different scenarios. We consider questions from PubMedQA and several tasks, ranging from binary (yes/no) responses to long answer generation, where the answer of the model is produced after an interaction with a physician. Our findings suggest that prompt design significantly influences the downstream accuracy of LLMs and that LLMs can provide valuable feedback to physicians, challenging incorrect diagnoses and contributing to more accurate decision-making. For example, when the physician is accurate 38% of the time, Mistral can produce the correct answer, improving accuracy up to 74% depending on the prompt being used, while Llama2 and Meditron models exhibit greater sensitivity to prompt choice. Our analysis also uncovers the challenges of ensuring that LLM-generated suggestions are pertinent and useful, emphasizing the need for further research in this area.

📄 PDF Abstract BibTeX arXiv:2403.20288

Code (1)

unitn-sml/physician-medLLM-interaction 공식 구현

Tasks

Answer GenerationDecision Making

Similar Papers 제목 키워드 기반

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

2026-06-17 · Tianming Du, Peijie Yu, Sihan Shang, Danli Shi 외 arxiv

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communicatio…

Clinical Knowledge

Improving Patient Pre-screening for Clinical Trials: Assisting Physicians with Large Language Models

2023-04-14 · Danny M. den Hamer, Perry Schoor, Tobias B. Polak, Daniel Kapitan

Physicians considering clinical trials for their patients are met with the laborious process of checking many text based eligibility criteria. Large Language Models (LLMs) have shown to perform well for clinical informat…

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

2025-05-29 · Yuexing Hao, Kumail Alhamoud, Hyewon Jeong, Haoran Zhang 외

Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logi…

Medical Question AnsweringQuestion AnsweringSentence

Assessing Empathy in Large Language Models with Real-World Physician-Patient Interactions

2024-05-26 · Man Luo, Christopher J. Warren, Lu Cheng, Haidar M. Abdul-Muhsin 외

The integration of Large Language Models (LLMs) into the healthcare domain has the potential to significantly enhance patient care and support through the development of empathetic, patient-facing chatbots. This study in…

LLM Sensitivity Evaluation Framework for Clinical Diagnosis

2025-04-18 · Chenwei Yan, Xiangling Fu, Yuxuan Xiong, Tianyi Wang 외

Large language models (LLMs) have demonstrated impressive performance across various domains. However, for clinical diagnosis, higher expectations are required for LLM's reliability and sensitivity: thinking like physici…

Decision MakingDiagnosticSensitivity