paper-with-me

홈 › Papers

Comparative Analysis of Large Language Models in Healthcare

2026-04-11 · Subin Santhosh, Farwa Abbas, Hussain Ahmad, Claudia Szabo arxiv

Background: Large Language Models (LLMs) are transforming artificial intelligence applications in healthcare due to their ability to understand, generate, and summarize complex medical text. They offer valuable support to clinicians, researchers, and patients, yet their deployment in high-stakes clinical environments raises critical concerns regarding accuracy, reliability, and patient safety. Despite substantial attention in recent years, standardized benchmarking of LLMs for medical applications has been limited. Objective: This study addresses the need for a standardized comparative evaluation of LLMs in medical settings. Method: We evaluate multiple models, including ChatGPT, LLaMA, Grok, Gemini, and ChatDoctor, on core medical tasks such as patient note summarization and medical question answering, using the open-access datasets, MedMCQA, PubMedQA, and Asclepius, and assess performance through a combination of linguistic and task-specific metrics. Results: The results indicate that domain-specific models, such as ChatDoctor, excel in contextual reliability, producing medically accurate and semantically aligned text, whereas general-purpose models like Grok and LLaMA perform better in structured question-answering tasks, demonstrating higher quantitative accuracy. This highlights the complementary strengths of domain-specific and general-purpose LLMs depending on the medical task. Conclusion: Our findings suggest that LLMs can meaningfully support medical professionals and enhance clinical decision-making; however, their safe and effective deployment requires adherence to ethical standards, contextual accuracy, and human oversight in relevant cases. These results underscore the importance of task-specific evaluation and cautious integration of LLMs into healthcare workflows.

📄 PDF Abstract BibTeX arXiv:2604.10316

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Comparative Analysis of Drug-GPT and ChatGPT LLMs for Healthcare Insights: Evaluating Accuracy and Relevance in Patient and HCP Contexts

2023-07-24 · Giorgos Lysandrou, Roma English Owen, Kirsty Mursec, Grant Le Brun 외

This study presents a comparative analysis of three Generative Pre-trained Transformer (GPT) solutions in a question and answer (Q&A) setting: Drug-GPT 3, Drug-GPT 4, and ChatGPT, in the context of healthcare application…

Understanding Stigmatizing Language Lexicons: A Comparative Analysis in Clinical Contexts

2025-09-09 · Yiliang Zhou, Di Hu, Tianchu Lyu, Jasmine Dhillon 외 arxiv

Stigmatizing language results in healthcare inequities, yet there is no universally accepted or standardized lexicon defining which words, terms, or phrases constitute stigmatizing language in healthcare. We conducted a …

Semantic SimilaritySentiment Analysis

Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation

2026-09-15 · Ting-Wei Chang, Hen-Hsen Huang, Hsin-Hsi Chen arxiv

This paper addresses the challenge of making complex healthcare information more accessible through automated Plain Language Adaptation (PLA). PLA aims to simplify technical medical language, bridging a critical gap betw…

Reading ComprehensionText SimplificationFew-Shot Learning

Are Large Language Models Ready for Healthcare? A Comparative Study on Clinical Language Understanding

2023-04-09 · Yuqing Wang, Yun Zhao, Linda Petzold

Large language models (LLMs) have made significant progress in various domains, including healthcare. However, the specialized nature of clinical language understanding tasks presents unique challenges and limitations th…

Document Classificationnamed-entity-recognitionNamed Entity RecognitionNatural Language Inference+4

Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis

2023-08-21 · Md. Arid Hasan, Shudipta Das, Afiyat Anjum, Firoj Alam 외

The rapid expansion of the digital world has propelled sentiment analysis into a critical tool across diverse sectors such as marketing, politics, customer service, and healthcare. While there have been significant advan…

In-Context LearningMarketingSentiment Analysis