paper-with-me

Papers

How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making?

2024-10-21 · Kenza Benkirane, Jackie Kay, Maria Perez-Ortiz

Recent advancements in Large Language Models (LLMs) have positioned them as powerful tools for clinical decision-making, with rapidly expanding applications in healthcare. However, concerns about bias remain a significant challenge in the clinical implementation of LLMs, particularly regarding gender and ethnicity. This research investigates the evaluation and mitigation of bias in LLMs applied to complex clinical cases, focusing on gender and ethnicity biases. We introduce a novel Counterfactual Patient Variations (CPV) dataset derived from the JAMA Clinical Challenge. Using this dataset, we built a framework for bias evaluation, employing both Multiple Choice Questions (MCQs) and corresponding explanations. We explore prompting with eight LLMs and fine-tuning as debiasing methods. Our findings reveal that addressing social biases in LLMs requires a multidimensional approach as mitigating gender bias can occur while introducing ethnicity biases, and that gender bias in LLM embeddings varies significantly across medical specialities. We demonstrate that evaluating both MCQ response and explanation processes is crucial, as correct responses can be based on biased \textit{reasoning}. We provide a framework for evaluating LLM bias in real-world clinical cases, offer insights into the complex nature of bias in these models, and present strategies for bias mitigation.

📄 PDF Abstract BibTeX arXiv:2410.16574

Code (1)

kenza-ily/diagnose_treat_bias_llm 공식 구현

Tasks

counterfactualDecision MakingMultiple-choice

Similar Papers 제목 키워드 기반

Evaluate underdiagnosis and overdiagnosis bias of deep learning model on primary open-angle glaucoma diagnosis in under-served patient populations

2023-01-26 · Mingquan Lin, Yuyun Xiao, BoJian Hou, Tingyi Wanyan 외

In the United States, primary open-angle glaucoma (POAG) is the leading cause of blindness, especially among African American and Hispanic individuals. Deep learning has been widely used to detect POAG using fundus image…

Deep Learning

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

2026-04-16 · Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor 외 arxiv

Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM Jury, composed of three frontier AI…

Then and Now: Quantifying the Longitudinal Validity of Self-Disclosed Depression Diagnoses

2022-06-22 · NAACL (CLPsych) 2022 7 · Keith Harrigian, Mark Dredze

Self-disclosed mental health diagnoses, which serve as ground truth annotations of mental health status in the absence of clinical measures, underpin the conclusions behind most computational studies of mental health lan…

Selection bias

Leveraging Large Language Models to Extract Information on Substance Use Disorder Severity from Clinical Notes: A Zero-shot Learning Approach

2024-03-18 · Maria Mahbub, Gregory M. Dams, Sudarshan Srinivasan, Caitlin Rizy 외

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal…

DiagnosticZero-Shot Learning

Detecting clinician implicit biases in diagnoses using proximal causal inference

2025-01-27 · Kara Liu, Russ Altman, Vasilis Syrgkanis

Clinical decisions to treat and diagnose patients are affected by implicit biases formed by racism, ableism, sexism, and other stereotypes. These biases reflect broader systemic discrimination in healthcare and risk marg…

AttributeCausal Inference