paper-with-me

Papers

NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models

2025-04-29 · Yi Zhou, Wenpeng Xing, Dezhang Kong, Changting Lin, Meng Han

Safety alignment in large language models (LLMs) is achieved through fine-tuning mechanisms that regulate neuron activations to suppress harmful content. In this work, we propose a novel approach to induce disalignment by identifying and modifying the neurons responsible for safety constraints. Our method consists of three key steps: Neuron Activation Analysis, where we examine activation patterns in response to harmful and harmless prompts to detect neurons that are critical for distinguishing between harmful and harmless inputs; Similarity-Based Neuron Identification, which systematically locates the neurons responsible for safe alignment; and Neuron Relearning for Safety Removal, where we fine-tune these selected neurons to restore the model's ability to generate previously restricted responses. Experimental results demonstrate that our method effectively removes safety constraints with minimal fine-tuning, highlighting a critical vulnerability in current alignment techniques. Our findings underscore the need for robust defenses against adversarial fine-tuning attacks on LLMs.

📄 PDF Abstract BibTeX arXiv:2504.21053

Code (0)

등록된 구현이 없습니다.

Tasks

Safety Alignment

Similar Papers 제목 키워드 기반

Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

2024-02-19 · Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu 외

Large language models (LLMs) undergo safety alignment to ensure safe conversations with humans. However, this paper introduces a training-free attack method capable of reversing safety alignment, converting the outcomes …

Language ModelingLanguage ModellingSafety Alignment

Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond

2025-02-07 · Chongyu Fan, Jinghan Jia, Yihua Zhang, Anil Ramakrishna 외

The LLM unlearning technique has recently been introduced to comply with data regulations and address the safety and ethical concerns of LLMs by removing the undesired data-model influence. However, state-of-the-art unle…

Large Language Models Relearn Removed Concepts

2024-01-03 · Michelle Lo, Shay B. Cohen, Fazl Barez

Advances in model editing through neuron pruning hold promise for removing undesirable concepts from large language models. However, it remains unclear whether models have the capacity to reacquire pruned concepts after …

Model Editing

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

2026-05-12 · Zeguan Xiao, Xuanzhe Xu, Yun Chen, Yong Wang 외 arxiv

Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critica…

Revisiting the Role of Relearning in Semantic Dementia

2025-03-05 · Devon Jarvis, Verena Klar, Richard Klein, Benjamin Rosman 외

Patients with semantic dementia (SD) present with remarkably consistent atrophy of neurons in the anterior temporal lobe and behavioural impairments, such as graded loss of category knowledge. While relearning of lost kn…