paper-with-me

홈 › Papers

Towards Leveraging Large Language Models for Automated Medical Q&A Evaluation

2024-09-03 · Jack Krolik, Herprit Mahal, Feroz Ahmad, Gaurav Trivedi, Bahador Saket

This paper explores the potential of using Large Language Models (LLMs) to automate the evaluation of responses in medical Question and Answer (Q\&A) systems, a crucial form of Natural Language Processing. Traditionally, human evaluation has been indispensable for assessing the quality of these responses. However, manual evaluation by medical professionals is time-consuming and costly. Our study examines whether LLMs can reliably replicate human evaluations by using questions derived from patient data, thereby saving valuable time for medical experts. While the findings suggest promising results, further research is needed to address more specific or complex questions that were beyond the scope of this initial investigation.

📄 PDF Abstract BibTeX arXiv:2409.01941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-as-a-Fuzzy-Judge: Fine-Tuning Large Language Models as a Clinical Evaluation Judge with Fuzzy Logic

2025-06-12 · Weibing Zheng, Laurah Turner, Jess Kropczynski, Murat Ozer 외

Clinical communication skills are critical in medical education, and practicing and assessing clinical communication skills on a scale is challenging. Although LLM-powered clinical scenario simulations have shown promise…

Large Language ModelPrompt Engineering

Assessing Automated Fact-Checking for Medical LLM Responses with Knowledge Graphs

2025-11-16 · Shasha Zhou, Mingyu Huang, Jack Cole, Charles Britton 외 arxiv

The recent proliferation of large language models (LLMs) holds the potential to revolutionize healthcare, with strong capabilities in diverse medical tasks. Yet, deploying LLMs in high-stakes healthcare settings requires…

Knowledge Graphs

MedGo: A Chinese Medical Large Language Model

2024-10-27 · HaiTao Zhang, Bo An

Large models are a hot research topic in the field of artificial intelligence. Leveraging their generative capabilities has the potential to enhance the level and quality of medical services. In response to the limitatio…

Language ModelingLanguage ModellingLarge Language ModelMedical Question Answering+2

BioACE: An Automated Framework for Biomedical Answer and Citation Evaluations

2026-02-04 · Deepak Gupta, Davis Bartels, Dina Demner-Fushman arxiv

With the increasing use of large language models (LLMs) for generating answers to biomedical questions, it is crucial to evaluate the quality of the generated answers and the references provided to support the facts in t…

Natural Language InferenceQuestion Answering

MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering

2026-03-15 · Shaowei Guan, Yu Zhai, Hin Chi Kwok, Jiawei Du 외 arxiv

Recent advances in Retrieval-Augmented Generation (RAG) have enabled large language models (LLMs) to ground outputs in clinical evidence. However, connecting LLMs with external databases introduces the risk of contextual…

Natural Language InferenceQuestion Answering