paper-with-me

홈 › Papers

Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions

2025-10-15 · Siying Liu, Shisheng Zhang, Indu Bala arxiv

Large language models (LLMs) are increasingly applied in biomedical domains, yet their reliability in drug-safety prediction remains underexplored. In this work, we investigate whether LLMs incorporate socio-demographic information into adverse event (AE) predictions, despite such attributes being clinically irrelevant. Using structured data from the United States Food and Drug Administration Adverse Event Reporting System (FAERS) and a persona-based evaluation framework, we assess two state-of-the-art models, ChatGPT-4o and Bio-Medical-Llama-3.8B, across diverse personas defined by education, marital status, employment, insurance, language, housing stability, and religion. We further evaluate performance across three user roles (general practitioner, specialist, patient) to reflect real-world deployment scenarios where commercial systems often differentiate access by user type. Our results reveal systematic disparities in AE prediction accuracy. Disadvantaged groups (e.g., low education, unstable housing) were frequently assigned higher predicted AE likelihoods than more privileged groups (e.g., postgraduate-educated, privately insured). Beyond outcome disparities, we identify two distinct modes of bias: explicit bias, where incorrect predictions directly reference persona attributes in reasoning traces, and implicit bias, where predictions are inconsistent, yet personas are not explicitly mentioned. These findings expose critical risks in applying LLMs to pharmacovigilance and highlight the urgent need for fairness-aware evaluation protocols and mitigation strategies before clinical deployment.

📄 PDF Abstract BibTeX arXiv:2510.13931

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal dosing of anti-cancer treatment under drug-induced plasticity

2024-12-20 · Einar Bjarki Gunnarsson, Benedikt Vilji Magnússon, Jasmine Foo

While cancer has traditionally been considered a genetic disease, mounting evidence indicates an important role for non-genetic (epigenetic) mechanisms. Common anti-cancer drugs have recently been observed to induce the …

Exploring Drug Safety Through Knowledge Graphs: Protein Kinase Inhibitors as a Case Study

2026-02-17 · David Jackson, Michael Gertz, Jürgen Hesser arxiv

Adverse Drug Reactions (ADRs) are a leading cause of morbidity and mortality. Existing prediction methods rely mainly on chemical similarity, machine learning on structured databases, or isolated target profiles, but oft…

Knowledge Graphs

Prediction of Drug-Induced TdP Risks Using Machine Learning and Rabbit Ventricular Wedge Assay

2022-01-14 · Nan Miles Xi, Dalong Patrick Huang

The evaluation of drug-induced Torsades de pointes (TdP) risks is crucial in drug safety assessment. In this study, we discuss machine learning approaches in the prediction of drug-induced TdP risks using preclinical dat…

BIG-bench Machine LearningPrediction

Multimodal AI predicts clinical outcomes of drug combinations from preclinical data

2025-03-04 · Yepeng Huang, Xiaorui Su, Varun Ullanat, Ivy Liang 외

Predicting clinical outcomes from preclinical data is essential for identifying safe and effective drug combinations. Current models rely on structural or target-based features to identify high-efficacy, low-toxicity dru…

Large Language Model

From Bench to Bedside: A Review of Clinical Trials in Drug Discovery and Development

2024-12-12 · Tianyang Wang, Ming Liu, Benji Peng, Xinyuan Song 외

Clinical trials are an indispensable part of the drug development process, bridging the gap between basic research and clinical application. During the development of new drugs, clinical trials are used not only to evalu…

Drug DiscoveryMarketing