paper-with-me

홈 › Papers

People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy

2024-08-11 · Shruthi Shekar, Pat Pataranutaporn, Chethan Sarabu, Guillermo A. Cecchi, Pattie Maes

This paper presents a comprehensive analysis of how AI-generated medical responses are perceived and evaluated by non-experts. A total of 300 participants gave evaluations for medical responses that were either written by a medical doctor on an online healthcare platform, or generated by a large language model and labeled by physicians as having high or low accuracy. Results showed that participants could not effectively distinguish between AI-generated and Doctors' responses and demonstrated a preference for AI-generated responses, rating High Accuracy AI-generated responses as significantly more valid, trustworthy, and complete/satisfactory. Low Accuracy AI-generated responses on average performed very similar to Doctors' responses, if not more. Participants not only found these low-accuracy AI-generated responses to be valid, trustworthy, and complete/satisfactory but also indicated a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result of the response provided. This problematic reaction was comparable if not more to the reaction they displayed towards doctors' responses. This increased trust placed on inaccurate or inappropriate AI-generated medical advice can lead to misdiagnosis and harmful consequences for individuals seeking help. Further, participants were more trusting of High Accuracy AI-generated responses when told they were given by a doctor and experts rated AI-generated responses significantly higher when the source of the response was unknown. Both experts and non-experts exhibited bias, finding AI-generated responses to be more thorough and accurate than Doctors' responses but still valuing the involvement of a Doctor in the delivery of their medical advice. Ensuring AI systems are implemented with medical professionals should be the future of using AI for the delivery of medical advice.

📄 PDF Abstract BibTeX arXiv:2408.15266

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language Modelvalid

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

People on Drugs: Credibility of User Statements in Health Communities

2017-05-06 · Subhabrata Mukherjee, Gerhard Weikum, Cristian Danescu-Niculescu-Mizil

Online health communities are a valuable source of information for patients and physicians. However, such user-generated resources are often plagued by inaccuracies and misinformation. In this work we propose a method fo…

Misinformation

On the Generation of Medical Dialogues for COVID-19

2020-05-11 · Wenmian Yang, Guangtao Zeng, Bowen Tan, Zeqian Ju 외

Under the pandemic of COVID-19, people experiencing COVID19-related symptoms or exposed to risk factors have a pressing need to consult doctors. Due to hospital closure, a lot of consulting services have been moved onlin…

Dialogue GenerationTransfer Learning

MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models

2024-09-29 · Vibhor Agarwal, Yiqiao Jin, Mohit Chandra, Munmun De Choudhury 외

The remarkable capabilities of large language models (LLMs) in language understanding and generation have not rendered them immune to hallucinations. LLMs can still generate plausible-sounding but factually incorrect or …

Hallucination

Citations and Trust in LLM Generated Responses

2025-01-02 · Yifan Ding, Matthew Facciani, Amrit Poudel, Ellen Joyce 외

Question answering systems are rapidly advancing, but their opaque nature may impact user trust. We explored trust through an anti-monitoring framework, where trust is predicted to be correlated with presence of citation…

ChatbotQuestion Answering

eTracer: Towards Traceable Text Generation via Claim-Level Grounding

2026-01-07 · Bohao Chu, Qianli Wang, Hendrik Damm, Hui Wang 외 arxiv

How can system-generated responses be efficiently verified, especially in the high-stakes biomedical domain? To address this challenge, we introduce eTracer, a plug-and-play framework that enables traceable text generati…

Text Generation