paper-with-me

홈 › Papers

First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows

2026-04-20 · Sihao Xing, Zaur Gouliev arxiv

Large language models (LLMs) are increasingly used in clinical settings, raising concerns about racial bias in both generated medical text and clinical reasoning. Existing studies have identified bias in medical LLMs, but many focus on single models and give less attention to mitigation. This study uses the EU AI Act as a governance lens to evaluate five widely used LLMs across two tasks, namely synthetic patient-case generation and differential diagnosis ranking. Using race-stratified epidemiological distributions in the United States and expert differential diagnosis lists as benchmarks, we apply structured prompt templates and a two-part evaluation design to examine implicit and explicit racial bias. All models deviated from observed racial distributions in the synthetic case generation task, with GPT-4.1 showing the smallest overall deviation. In the differential diagnosis task, DeepSeek V3 produced the strongest overall results across the reported metrics. When embedded in an agentic workflow, DeepSeek V3 showed an improvement of 0.0348 in mean p-value, 0.1166 in median p-value, and 0.0949 in mean difference relative to the standalone model, although improvement was not uniform across every metric. These findings support multi-metric bias evaluation for AI systems used in medical settings and suggest that retrieval-based agentic workflows may reduce some forms of explicit bias in benchmarked diagnostic tasks. Detailed prompt templates, experimental datasets, and code pipelines are available on our GitHub.

📄 PDF Abstract BibTeX arXiv:2604.18038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Bias in Large Language Models: Fine-tuned KcBERT

2024-03-16 · J. K. Lee, T. M. Chung

The rapid advancement of large language models (LLMs) has enabled natural language processing capabilities similar to those of humans, and LLMs are being widely utilized across various societal domains such as education …

Language ModelingLanguage ModellingMasked Language Modeling

Bias Neutralization Framework: Measuring Fairness in Large Language Models with Bias Intelligence Quotient (BiQ)

2024-04-28 · Malur Narayan, John Pasmore, Elton Sampaio, Vijay Raghavan 외

The burgeoning influence of Large Language Models (LLMs) in shaping public discourse and decision-making underscores the imperative to address inherent biases within these AI systems. In the wake of AI's expansive integr…

Decision MakingFairnessLanguage ModelingLanguage Modelling+1

A Comprehensive Study of Implicit and Explicit Biases in Large Language Models

2025-11-18 · Fatima Kazi, Alex Young, Yash Inani, Setareh Rafatirad arxiv

Large Language Models (LLMs) inherit explicit and implicit biases from their training datasets. Identifying and mitigating biases in LLMs is crucial to ensure fair outputs, as they can perpetuate harmful stereotypes and …

Data Augmentation

Gender and Racial Stereotype Detection in Legal Opinion Word Embeddings

2022-03-24 · Sean Matthews, John Hudzina, Dawn Sepehr

Studies have shown that some Natural Language Processing (NLP) systems encode and replicate harmful biases with potential adverse ethical effects in our society. In this article, we propose an approach for identifying ge…

Question AnsweringWord Embeddings

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness

2025-12-22 · Naseem Machlovi, Maryam Saleki, Ruhul Amin, Mohamed Rahouti 외 arxiv

As large language models (LLMs) become deeply embedded in daily life, the urgent need for safer moderation systems that distinguish between naive and harmful requests while upholding appropriate censorship boundaries has…

Adversarial Robustness