paper-with-me

Papers

Appraising the Potential Uses and Harms of LLMs for Medical Systematic Reviews

2023-05-19 · Hye Sun Yun, Iain J. Marshall, Thomas A. Trikalinos, Byron C. Wallace

Medical systematic reviews play a vital role in healthcare decision making and policy. However, their production is time-consuming, limiting the availability of high-quality and up-to-date evidence summaries. Recent advancements in large language models (LLMs) offer the potential to automatically generate literature reviews on demand, addressing this issue. However, LLMs sometimes generate inaccurate (and potentially misleading) texts by hallucination or omission. In healthcare, this can make LLMs unusable at best and dangerous at worst. We conducted 16 interviews with international systematic review experts to characterize the perceived utility and risks of LLMs in the specific context of medical evidence reviews. Experts indicated that LLMs can assist in the writing process by drafting summaries, generating templates, distilling information, and crosschecking information. They also raised concerns regarding confidently composed but inaccurate LLM outputs and other potential downstream harms, including decreased accountability and proliferation of low-quality reviews. Informed by this qualitative analysis, we identify criteria for rigorous evaluation of biomedical LLMs aligned with domain expert views.

📄 PDF Abstract BibTeX arXiv:2305.11828

Code (1)

hyesunyun/medsysreviewsfromllms 공식 구현

Tasks

Decision MakingHallucination

Similar Papers 제목 키워드 기반

Overview of TREC 2024 Biomedical Generative Retrieval (BioGen) Track

2024-11-27 · Deepak Gupta, Dina Demner-Fushman, William Hersh, Steven Bedrick 외

With the advancement of large language models (LLMs), the biomedical domain has seen significant progress and improvement in multiple tasks such as biomedical question answering, lay language summarization of the biomedi…

Medical Question AnsweringQuestion AnsweringRetrieval

Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity

2026-01-31 · Prakhar Ganesh, Reza Shokri, Golnoosh Farnadi arxiv

Large language models (LLMs) are known to "hallucinate" by generating false or misleading outputs. Hallucinations pose various harms, from erosion of trust to widespread misinformation. Existing hallucination evaluation,…

AI Chatbots for Mental Health: Values and Harms from Lived Experiences of Depression

2025-04-26 · Dong Whi Yoo, Jiayue Melissa Shi, Violeta J. Rodriguez, Koustuv Saha

Recent advancements in LLMs enable chatbots to interact with individuals on a range of queries, including sensitive mental health contexts. Despite uncertainties about their effectiveness and reliability, the development…

ChatbotManagement

Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models

2023-09-15 · Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk 외

Large Language Models (LLMs), now used daily by millions, can encode societal biases, exposing their users to representational harms. A large body of scholarship on LLM bias exists but it predominantly adopts a Western-c…

FairnessLanguage ModellingLarge Language Model

mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

2025-09-02 · Shreyash Adappanavar, Krithi Shailya, Gokul S Krishnan, Sriraam Natarajan 외 arxiv

The deployment of Large Language Models (LLMs) in high-stakes medical settings poses a critical AI alignment challenge, as models can inherit and amplify societal biases, leading to significant disparities. Existing fair…