paper-with-me

홈 › Papers

Can SAEs reveal and mitigate racial biases of LLMs in healthcare?

2025-10-31 · Hiba Ahsan, Byron C. Wallace arxiv

LLMs are increasingly being used in healthcare. This promises to free physicians from drudgery, enabling better care to be delivered at scale. But the use of LLMs in this space also brings risks; for example, such models may worsen existing biases. How can we spot when LLMs are (spuriously) relying on patient race to inform predictions? In this work we assess the degree to which Sparse Autoencoders (SAEs) can reveal (and control) associations the model has made between race and stigmatizing concepts. We first identify SAE latents in Gemma-2 models which appear to correlate with Black individuals. We find that this latent activates on reasonable input sequences (e.g., "African American") but also problematic words like "incarceration". We then show that we can use this latent to steer models to generate outputs about Black patients, and further that this can induce problematic associations in model outputs as a result. For example, activating the Black latent increases the risk assigned to the probability that a patient will become "belligerent". We evaluate the degree to which such steering via latents might be useful for mitigating bias. We find that this offers improvements in simple settings, but is less successful for more realistic and complex clinical tasks. Overall, our results suggest that: SAEs may offer a useful tool in clinical applications of LLMs to identify problematic reliance on demographics but mitigating bias via SAE steering appears to be of marginal utility for realistic tasks.

📄 PDF Abstract BibTeX arXiv:2511.00177

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ChatGPT Exhibits Gender and Racial Biases in Acute Coronary Syndrome Management

2023-11-10 · Angela Zhang, Mert Yuksekgonul, Joshua Guild, James Zou 외

Recent breakthroughs in large language models (LLMs) have led to their rapid dissemination and widespread use. One early application has been to medicine, where LLMs have been investigated to streamline clinical workflow…

Decision MakingManagement

Evaluating Gender, Racial, and Age Biases in Large Language Models: A Comparative Analysis of Occupational and Crime Scenarios

2024-09-22 · Vishal Mirza, Rahul Kulkarni, Aakanksha Jadhav

Recent advancements in Large Language Models(LLMs) have been notable, yet widespread enterprise adoption remains limited due to various constraints. This paper examines bias in LLMs-a crucial issue affecting their usabil…

Fairness

Bias Neutralization Framework: Measuring Fairness in Large Language Models with Bias Intelligence Quotient (BiQ)

2024-04-28 · Malur Narayan, John Pasmore, Elton Sampaio, Vijay Raghavan 외

The burgeoning influence of Large Language Models (LLMs) in shaping public discourse and decision-making underscores the imperative to address inherent biases within these AI systems. In the wake of AI's expansive integr…

Decision MakingFairnessLanguage ModelingLanguage Modelling+1

BLIND: Bias Removal With No Demographics

2022-12-20 · Hadas Orgad, Yonatan Belinkov

Models trained on real-world data tend to imitate and amplify social biases. Common methods to mitigate biases require prior information on the types of biases that should be mitigated (e.g., gender or racial bias) and t…

Sentiment AnalysisSentiment Classification

Evaluation of Bias Towards Medical Professionals in Large Language Models

2024-06-30 · Xi Chen, Yang Xu, MingKe You, Li Wang 외

This study evaluates whether large language models (LLMs) exhibit biases towards medical professionals. Fictitious candidate resumes were created to control for identity factors while maintaining consistent qualification…