paper-with-me

Papers

Overriding Safety protections of Open-source Models

2024-09-28 · Sachin Kumar

LLMs(Large Language Models) nowadays have widespread adoption as a tool for solving issues across various domain/tasks. These models since are susceptible to produce harmful or toxic results, inference-time adversarial attacks, therefore they do undergo safety alignment training and Red teaming for putting in safety guardrails. For using these models, usually fine-tuning is done for model alignment on the desired tasks, which can make model more aligned but also make it more susceptible to produce unsafe responses, if fine-tuned with harmful data.In this paper, we study how much of impact introduction of harmful data in fine-tuning can make, and if it can override the safety protection of those models. Conversely,it was also explored that if model is fine-tuned on safety data can make the model produce more safer responses. Further we explore if fine-tuning the model on harmful data makes it less helpful or less trustworthy because of increase in model uncertainty leading to knowledge drift. Our extensive experimental results shown that Safety protection in an open-source can be overridden, when fine-tuned with harmful data as observed by ASR increasing by 35% when compared to basemodel's ASR. Also, as observed, fine-tuning a model with harmful data made the harmful fine-tuned model highly uncertain with huge knowledge drift and less truthfulness in its responses. Furthermore, for the safe fine-tuned model, ASR decreases by 51.68% as compared to the basemodel, and Safe model also shown in minor drop in uncertainty and truthfulness as compared to basemodel. This paper's code is available at: https://github.com/techsachinkr/Overriding_Model_Safety_Protections

📄 PDF Abstract BibTeX arXiv:2409.19476

Code (1)

techsachinkr/Overriding_Model_Safety_Protections 공식 구현

Tasks

Red TeamingSafety Alignment

Similar Papers 제목 키워드 기반

DiffGuard: Text-Based Safety Checker for Diffusion Models

2024-11-25 · Massine El Khader, Elias Al Bouzidi, Abdellah Oumida, Mohammed Sbaihi 외

Recent advances in Diffusion Models have enabled the generation of images from text, with powerful closed-source models like DALL-E and Midjourney leading the way. However, open-source alternatives, such as StabilityAI's…

Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety

2025-10-13 · Wei-Chieh Huang, Henry Peng Zou, Yaozu Wu, Dongyuan Li 외 arxiv

Deep research frameworks have shown promising capabilities in synthesizing comprehensive reports from web sources. While deep research possesses significant potential to address complex issues through planning and resear…

Gameplay Filters: Robust Zero-Shot Safety through Adversarial Imagination

2024-05-01 · Duy P. Nguyen, Kai-Chieh Hsu, Wenhao Yu, Jie Tan 외

Despite the impressive recent advances in learning-based robot control, ensuring robustness to out-of-distribution conditions remains an open challenge. Safety filters can, in principle, keep arbitrary control policies f…

AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?

2025-10-13 · Leonard Dung, Florian Mai arxiv

AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the …

Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

2026-05-20 · Roland Pihlakas, Jan Llenzl Dagohoy arxiv

Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behaviour of LLMs under sustained authority pre…