paper-with-me

홈 › Papers

Red Teaming Large Language Models for Healthcare

2025-05-01 · Vahid Balazadeh, Michael Cooper, David Pellow, Atousa Assadi, Jennifer Bell, Jim Fackler, Gabriel Funingana, Spencer Gable-Cook, Anirudh Gangadhar, Abhishek Jaiswal, Sumanth Kaja, Christopher Khoury, Randy Lin, Kaden McKeen, Sara Naimimohasses, Khashayar Namdar, Aviraj Newatia, Allan Pang, Anshul Pattoo, Sameer Peesapati, Diana Prepelita, Bogdana Rakova, Saba Sadatamin, Rafael Schulman, Ajay Shah, Syed Azhar Shah, Syed Ahmar Shah, Babak Taati, Balagopal Unnikrishnan, Stephanie Williams, Rahul G Krishnan

We present the design process and findings of the pre-conference workshop at the Machine Learning for Healthcare Conference (2024) entitled Red Teaming Large Language Models for Healthcare, which took place on August 15, 2024. Conference participants, comprising a mix of computational and clinical expertise, attempted to discover vulnerabilities -- realistic clinical prompts for which a large language model (LLM) outputs a response that could cause clinical harm. Red-teaming with clinicians enables the identification of LLM vulnerabilities that may not be recognised by LLM developers lacking clinical expertise. We report the vulnerabilities found, categorise them, and present the results of a replication study assessing the vulnerabilities across all LLMs provided.

📄 PDF Abstract BibTeX arXiv:2505.00467

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelRed Teaming

Similar Papers 제목 키워드 기반

Aloe: A Family of Fine-tuned Open Healthcare LLMs

2024-05-03 · Ashwin Kumar Gururajan, Enrique Lopez-Cuena, Jordi Bayarri-Planas, Adrian Tormos 외

As the capabilities of Large Language Models (LLMs) in healthcare and medicine continue to advance, there is a growing need for competitive open-source models that can safeguard public interest. With the increasing avail…

Prompt EngineeringRed Teaming

The case for delegated AI autonomy for Human AI teaming in healthcare

2025-03-24 · Yan Jia, Harriet Evans, Zoe Porter, Simon Graham 외

In this paper we propose an advanced approach to integrating artificial intelligence (AI) into healthcare: autonomous decision support. This approach allows the AI algorithm to act autonomously for a subset of patient ca…

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

2026-05-05 · Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers arxiv

AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current approaches force operators into manual, lib…

Red Teaming

Advancing Human-Machine Teaming: Concepts, Challenges, and Applications

2025-03-16 · Dian Chen, Han Jun Yoon, Zelin Wan, Nithin Alluru 외

Human-Machine Teaming (HMT) is revolutionizing collaboration across domains such as defense, healthcare, and autonomous systems by integrating AI-driven decision-making, trust calibration, and adaptive teaming. This surv…

BenchmarkingDecision MakingDomain Adaptation

Do No Harm: Exposing Hidden Vulnerabilities of LLMs via Persona-based Client Simulation Attack in Psychological Counseling

2026-04-06 · Qingyang Xu, Yaling Shen, Stephanie Fong, Zimu Wang 외 arxiv

The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions. A key challenge is distinguishing therapeutic empathy from maladaptive validation, w…