paper-with-me

Papers

Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency

2026-06-02 · Qi Han Wong arxiv

We investigate whether large language models produce different medical triage recommendations for identical neurological symptoms when only the patient's stated gender and age vary. Using three model families--Gemini 3.5 Flash, Claude Sonnet 4.6, and GPT-5.4-mini--we present a standardized symptom profile (persistent headache, blurred vision, morning nausea, visual disturbances) across seven demographic conditions: three age groups (25, 38, 65) x two genders (male, female), plus a gender-unspecified baseline (n = 30 per condition per model, 630 total trials). We find a stark, systemic gender-dependent triage disparity: young women receive significantly lower emergency room (ER) referral rates than age-matched men (Gemini: 0% vs. 23.3%; Claude: 6.7% vs. 96.7%; GPT: 6.7% vs. 66.7%, all p < 0.001). The disparity disappears at age 65 for all models. The primary mechanism is diagnostic substitution: the models anchor on a gender-associated diagnosis, preferentially classifying young women with Idiopathic Intracranial Hypertension (IIH)--a condition epidemiologically linked to women of childbearing age--while diagnosing men with generic increased intracranial pressure with space-occupying lesions in the differential. This diagnostic closure routes female patients to lower-urgency care (outpatient doctor appointments) despite comparable severity ratings (7-9/10). Our findings demonstrate that clinical LLMs replicate documented human clinical biases by using epidemiological priors to suppress triage urgency, suggesting that AI triage engines must decouple urgency assessment from probabilistic diagnostic priors. We release all code, prompts, and raw results.

📄 PDF Abstract BibTeX arXiv:2606.03641

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A comparative study of artificial intelligence and human doctors for the purpose of triage and diagnosis

2018-06-27 · Salman Razzaki, Adam Baker, Yura Perov, Katherine Middleton 외

Online symptom checkers have significant potential to improve patient care, however their reliability and accuracy remain variable. We hypothesised that an artificial intelligence (AI) powered triage and diagnostic syste…

Diagnostic

EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage

2026-05-05 · Richard J. Young, Alice M. Matthews arxiv

Emergency department triage assigns patients an acuity score that determines treatment priority, and clinical evidence documents persistent gender disparities in human acuity assessment. As hospitals pilot large language…

Superhuman performance of a large language model on the reasoning tasks of a physician

2024-12-14 · Peter G. Brodeur, Thomas A. Buckley, Zahir Kanjee, Ethan Goh 외

A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever sin…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model+2

MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage

2026-03-24 · Ufaq Khan, Umair Nawaz, L D M S S Teja, Numaan Saeed 외 arxiv

Vision Language Models (VLMs) are increasingly used for tasks like medical report generation and visual question answering. However, fluent diagnostic text does not guarantee safe visual understanding. In clinical practi…

Visual Question AnsweringMedical Report Generation

Nichelle and Nancy: The Influence of Demographic Attributes and Tokenization Length on First Name Biases

2023-05-26 · Haozhe An, Rachel Rudinger

Through the use of first name substitution experiments, prior research has demonstrated the tendency of social commonsense reasoning models to systematically exhibit social biases along the dimensions of race, ethnicity,…