paper-with-me

홈 › Papers

The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness

2023-12-30 · Neeraj Varshney, Pavel Dolin, Agastya Seth, Chitta Baral

As Large Language Models (LLMs) play an increasingly pivotal role in natural language processing applications, their safety concerns become critical areas of NLP research. This paper presents Safety and Over-Defensiveness Evaluation (SODE) benchmark: a collection of diverse safe and unsafe prompts with carefully designed evaluation methods that facilitate systematic evaluation, comparison, and analysis over 'safety' and 'over-defensiveness.' With SODE, we study a variety of LLM defense strategies over multiple state-of-the-art LLMs, which reveals several interesting and important findings, such as (a) the widely popular 'self-checking' techniques indeed improve the safety against unsafe inputs, but this comes at the cost of extreme over-defensiveness on the safe inputs, (b) providing a safety instruction along with in-context exemplars (of both safe and unsafe inputs) consistently improves safety and also mitigates undue over-defensiveness of the models, (c) providing contextual knowledge easily breaks the safety guardrails and makes the models more vulnerable to generating unsafe responses. Overall, our work reveals numerous such critical findings that we believe will pave the way and facilitate further research in improving the safety of LLMs.

📄 PDF Abstract BibTeX arXiv:2401.00287

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Data Reconstruction Attacks and Defenses: A Systematic Evaluation

2024-02-13 · Sheng Liu, Zihan Wang, Yuxiao Chen, Qi Lei

Reconstruction attacks and defenses are essential in understanding the data leakage problem in machine learning. However, prior work has centered around empirical observations of gradient inversion attacks, lacks theoret…

Federated LearningReconstruction Attack

Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations

2025-11-24 · Ryan Wong, Hosea David Yu Fei Ng, Dhananjai Sharma, Glenn Jun Jie Ng 외 arxiv

Large Language Models (LLMs) remain susceptible to jailbreak exploits that bypass safety filters and induce harmful or unethical behavior. This work presents a systematic taxonomy of existing jailbreak defenses across pr…

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

2024-11-13 · Yangyang Guo, Fangkai Jiao, Liqiang Nie, Mohan Kankanhalli

The vulnerability of Vision Large Language Models (VLLMs) to jailbreak attacks appears as no surprise. However, recent defense mechanisms against these attacks have reached near-saturation performance on benchmark evalua…

Towards Attack-tolerant Federated Learning via Critical Parameter Analysis

2023-08-18 · ICCV 2023 1 · Sungwon Han, Sungwon Park, Fangzhao Wu, Sundong Kim 외

Federated learning is used to train a shared model in a decentralized way without clients sharing private data with each other. Federated learning systems are susceptible to poisoning attacks when malicious clients send …

Federated Learning

Design and Analysis of Novel Bit-flip Attacks and Defense Strategies for DNNs

2022-06-24 · IEEE Conference on Dependable and Secure Computing (DSC) 2022 6 · Yash Khare, Kumud Lakara, Maruthi S Inukonda, Sparsh Mittal 외

The security of deep neural networks (DNNs) has become a matter of grave concern in the past few years due to their increasing ubiquity in security-critical domains. In this paper, we present novel bit-flip attack (BFA) …