How Reliable AI Chatbots are for Disease Prediction from Patient Complaints?
Artificial Intelligence (AI) chatbots leveraging Large Language Models (LLMs) are gaining traction in healthcare for their potential to automate patient interactions and aid clinical decision-making. This study examines the reliability of AI chatbots, specifically GPT 4.0, Claude 3 Opus, and Gemini Ultra 1.0, in predicting diseases from patient complaints in the emergency department. The methodology includes few-shot learning techniques to evaluate the chatbots' effectiveness in disease prediction. We also fine-tune the transformer-based model BERT and compare its performance with the AI chatbots. Results suggest that GPT 4.0 achieves high accuracy with increased few-shot data, while Gemini Ultra 1.0 performs well with fewer examples, and Claude 3 Opus maintains consistent performance. BERT's performance, however, is lower than all the chatbots, indicating limitations due to limited labeled data. Despite the chatbots' varying accuracy, none of them are sufficiently reliable for critical medical decision-making, underscoring the need for rigorous validation and human oversight. This study reflects that while AI chatbots have potential in healthcare, they should complement, not replace, human expertise to ensure patient safety. Further refinement and research are needed to improve AI-based healthcare applications' reliability for disease prediction.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingDisease PredictionFew-Shot LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The Relationship Between Head Injury and Alzheimer's Disease: A Causal Analysis with Bayesian Networks
This study examines the potential causal relationship between head injury and the risk of developing Alzheimer's disease (AD) using Bayesian networks and regression models. Using a dataset of 2,149 patients, we analyze k…
regressionDistinguishing Scams and Fraud with Ensemble Learning
Users increasingly query LLM-enabled web chatbots for help with scam defense. The Consumer Financial Protection Bureau's complaints database is a rich data source for evaluating LLM performance on user scam queries, but …
Ensemble LearningWearable-based Mediation State Detection in Individuals with Parkinson's Disease
One of the most prevalent complaints of individuals with mid-stage and advanced Parkinson's disease (PD) is the fluctuating response to their medication (i.e., ON state with maximum benefit from medication and OFF state …
SpecificityAuditory verbal learning disabilities in patients with mild cog impairment and mild Alzheimer's disease: A clinical study
Learning and memory impairments are common characteristics of individuals with mild cognitive impairment (MCI) and mild Alzheimer's disease (miAD). Early diagnosis of MCI is necessary to prevent recurrence of the disease…
Medical Image Super-Resolution Using a Generative Adversarial Network
During the growing popularity of electronic medical records, electronic medical record (EMR) data has exploded increasingly. It is very meaningful to retrieve high quality EMR in mass data. In this paper, an EMR value ne…
Brain SegmentationContent-Based Image RetrievalGenerative Adversarial NetworkImage Generation+4