paper-with-me

Papers

CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs

2025-05-16 · Sijia Chen, Xiaomin Li, Mengxue Zhang, Eric Hanchen Jiang, Qingcheng Zeng, Chen-Hsiang Yu

Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded harmfulness levels, and coverage of jailbreak-style attacks. We introduce CARES (Clinical Adversarial Robustness and Evaluation of Safety), a benchmark for evaluating LLM safety in healthcare. CARES includes over 18,000 prompts spanning eight medical safety principles, four harm levels, and four prompting styles: direct, indirect, obfuscated, and role-play, to simulate both malicious and benign use cases. We propose a three-way response evaluation protocol (Accept, Caution, Refuse) and a fine-grained Safety Score metric to assess model behavior. Our analysis reveals that many state-of-the-art LLMs remain vulnerable to jailbreaks that subtly rephrase harmful prompts, while also over-refusing safe but atypically phrased queries. Finally, we propose a mitigation strategy using a lightweight classifier to detect jailbreak attempts and steer models toward safer behavior via reminder-based conditioning. CARES provides a rigorous framework for testing and improving medical LLM safety under adversarial and ambiguous conditions.

📄 PDF Abstract BibTeX arXiv:2505.11413

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessSafety AlignmentSpecificity

Similar Papers 제목 키워드 기반

CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

2024-06-10 · Peng Xia, Ze Chen, Juanxi Tian, Yangrui Gong 외

Artificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized he…

Fairness

Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment

2025-12-03 · Huy Nghiem, Swetasudha Panda, Devashish Khatwani, Huy V. Nguyen 외 arxiv

Large Language Models (LLMs) are increasingly used in healthcare, yet ensuring their safety and trustworthiness remains a barrier to deployment. Conversational medical assistants must avoid unsafe compliance without over…

Adversarial Robustness

Robustness of Large Language Models Against Adversarial Attacks

2024-12-22 · Yiyi Tao, Yixian Shen, Hang Zhang, Yanxin Shen 외

The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the…

Sentiment AnalysisSentiment ClassificationSST-2

ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

2023-10-14 · Alex Mei, Sharon Levy, William Yang Wang

As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustness evaluations must comprehensively enca…

Red Teaming

Probabilistic Robustness in Deep Learning: A Concise yet Comprehensive Guide

2025-02-20 · Xingyu Zhao

Deep learning (DL) has demonstrated significant potential across various safety-critical applications, yet ensuring its robustness remains a key challenge. While adversarial robustness has been extensively studied in wor…

Adversarial RobustnessBenchmarkingDeep Learning