paper-with-me

홈 › Papers

LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation

2023-08-15 · Xiaoming Shi, Jie Xu, Jinru Ding, Jiali Pang, Sichen Liu, Shuqing Luo, Xingwei Peng, Lu Lu, Haihong Yang, Mingtao Hu, Tong Ruan, Shaoting Zhang

There is an increasing interest in developing LLMs for medical diagnosis to improve diagnosis efficiency. Despite their alluring technological potential, there is no unified and comprehensive evaluation criterion, leading to the inability to evaluate the quality and potential risks of medical LLMs, further hindering the application of LLMs in medical treatment scenarios. Besides, current evaluations heavily rely on labor-intensive interactions with LLMs to obtain diagnostic dialogues and human evaluation on the quality of diagnosis dialogue. To tackle the lack of unified and comprehensive evaluation criterion, we first initially establish an evaluation criterion, termed LLM-specific Mini-CEX to assess the diagnostic capabilities of LLMs effectively, based on original Mini-CEX. To address the labor-intensive interaction problem, we develop a patient simulator to engage in automatic conversations with LLMs, and utilize ChatGPT for evaluating diagnosis dialogues automatically. Experimental results show that the LLM-specific Mini-CEX is adequate and necessary to evaluate medical diagnosis dialogue. Besides, ChatGPT can replace manual evaluation on the metrics of humanistic qualities and provides reproducible and automated comparisons between different LLMs.

📄 PDF Abstract BibTeX arXiv:2308.07635

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticLanguage ModelingLanguage ModellingLarge Language ModelMedical Diagnosis

Similar Papers 제목 키워드 기반

CephGPT-4: An Interactive Multimodal Cephalometric Measurement and Diagnostic System with Visual Large Language Model

2023-07-01 · Lei Ma, Jincong Han, Zhaoxin Wang, Dian Zhang

Large-scale multimodal language models (LMMs) have achieved remarkable success in general domains. However, the exploration of diagnostic language models based on multimodal cephalometric medical data remains limited. In…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model

Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries

2024-09-19 · Kiran Vodrahalli, Santiago Ontanon, Nilesh Tripuraneni, Kelvin Xu 외

We introduce Michelangelo: a minimal, synthetic, and unleaked long-context reasoning evaluation for large language models which is also easy to automatically score. This evaluation is derived via a novel, unifying framew…

DiagnosticLanguage ModelingLanguage Modelling

Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation

2026-01-20 · Arjun Chandra, Kevin Miller, Venkatesh Ravichandran, Constantinos Papayiannis 외 arxiv

Large Language Model (LLM) judges exhibit strong reasoning capabilities but are limited to textual content. This leaves current automatic Speech-to-Speech (S2S) evaluation methods reliant on opaque and expensive Audio La…

SymptomWise: A Deterministic Reasoning Layer for Reliable and Efficient AI Systems

2026-04-07 · Isaac Henry, Avery Byrne, Christopher Giza, Ron Henry 외 arxiv

AI-driven symptom analysis systems face persistent challenges in reliability, interpretability, and hallucination. End-to-end generative approaches often lack traceability and may produce unsupported or inconsistent diag…

A Data-Driven Guided Decoding Mechanism for Diagnostic Captioning

2024-06-20 · Panagiotis Kaliosis, John Pavlopoulos, Foivos Charalampakos, Georgios Moschovis 외

Diagnostic Captioning (DC) automatically generates a diagnostic text from one or more medical images (e.g., X-rays, MRIs) of a patient. Treated as a draft, the generated text may assist clinicians, by providing an initia…

DiagnosticImage to textText GenerationZero-Shot Learning