paper-with-me

홈 › Papers

Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization

2025-10-01 · Sarvesh Soni, Dina Demner-Fushman arxiv

Automated approaches to answer patient-posed health questions are rising, but selecting among systems requires reliable evaluation. The current gold standard for evaluating the free-text artificial intelligence (AI) responses--human expert review--is labor-intensive and slow, limiting scalability. Automated metrics are promising yet variably aligned with human judgments and often context-dependent. To address the feasibility of automating the evaluation of AI responses to hospitalization-related questions posed by patients, we conducted a large systematic study of evaluation approaches. Across 100 patient cases, we collected responses from 28 AI systems (2800 total) and assessed them along three dimensions: whether a system response (1) answers the question, (2) appropriately uses clinical note evidence, and (3) uses general medical knowledge. Using clinician-authored reference answers to anchor metrics, automated rankings closely matched human ratings. Our findings suggest that carefully designed automated evaluation can scale comparative assessment of AI systems and support patient-clinician communication.

📄 PDF Abstract BibTeX arXiv:2510.00436

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Scalable Framework for Evaluating Health Language Models

2025-03-30 · Neil Mallinar, A. Ali Heydari, Xin Liu, Anthony Z. Faranesh 외

Large language models (LLMs) have emerged as powerful tools for analyzing complex datasets. Recent studies demonstrate their potential to generate useful, personalized responses when provided with patient-specific health…

Towards Leveraging Large Language Models for Automated Medical Q&A Evaluation

2024-09-03 · Jack Krolik, Herprit Mahal, Feroz Ahmad, Gaurav Trivedi 외

This paper explores the potential of using Large Language Models (LLMs) to automate the evaluation of responses in medical Question and Answer (Q\&A) systems, a crucial form of Natural Language Processing. Traditionally,…

Assessing Automated Fact-Checking for Medical LLM Responses with Knowledge Graphs

2025-11-16 · Shasha Zhou, Mingyu Huang, Jack Cole, Charles Britton 외 arxiv

The recent proliferation of large language models (LLMs) holds the potential to revolutionize healthcare, with strong capabilities in diverse medical tasks. Yet, deploying LLMs in high-stakes healthcare settings requires…

Knowledge Graphs

A Method to Automate the Discharge Summary Hospital Course for Neurology Patients

2023-05-10 · Vince C. Hartman, Sanika S. Bapat, Mark G. Weiner, Babak B. Navi 외

Generation of automated clinical notes have been posited as a strategy to mitigate physician burnout. In particular, an automated narrative summary of a patient's hospital stay could supplement the hospital course sectio…

Decoder

Assessing Empathy in Large Language Models with Real-World Physician-Patient Interactions

2024-05-26 · Man Luo, Christopher J. Warren, Lu Cheng, Haidar M. Abdul-Muhsin 외

The integration of Large Language Models (LLMs) into the healthcare domain has the potential to significantly enhance patient care and support through the development of empathetic, patient-facing chatbots. This study in…