paper-with-me

홈 › Papers

From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs

2025-11-18 · Erum Mushtaq, Anil Ramakrishna, Satyapriya Krishna, Sattvik Sahai, Prasoon Goyal, Kai-Wei Chang, Tao Zhang, Rahul Gupta arxiv

Recent work has shown that fine-tuning on insecure code data can trigger an emergent misalignment (EMA) phenomenon, where models generate malicious responses even to prompts unrelated to the original insecure code-writing task. Such cross-domain generalization of harmful behavior underscores the need for a deeper understanding of the algorithms, tasks, and datasets that induce emergent misalignment. In this work, we extend this study by demonstrating that emergent misalignment can also arise from narrow refusal unlearning in specific domains. We perform refusal unlearning on Cybersecurity and Safety concept, and evaluate EMA by monitoring refusal scores across seven responsible AI (RAI) domains, Cybersecurity, Safety, Toxicity, Bias, Sensitive Content, Medical/Legal, and Privacy. Our work shows that narrow domain unlearning can yield compliance responses for the targeted concept, however, it may also propagate EMA to unrelated domains. Among the two intervened concepts, Cybersecurity and Safety, we find that the safety concept can have larger EMA impact, i.e, causing lower refusal scores, across other unrelated domains such as bias. We observe this effect consistently across two model families, Mistral-7b-0.3v, and Qwen-7b-2.5. Further, we show that refusal unlearning augmented with cross-entropy loss function on a small set of retain data from the affected domains can largely, if not fully, restore alignment across the impacted domains while having lower refusal rate on the concept we perform unlearning on. To investigate the underlying causes of EMA, we analyze concept entanglements at the representation level via concept vectors. Our analysis reveals that concepts with higher representation similarity in earlier layers are more susceptible to EMA after intervention when the refusal stream is altered through targeted refusal unlearning.

📄 PDF Abstract BibTeX arXiv:2511.14017

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Similar Papers 제목 키워드 기반

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

2025-02-24 · Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley 외

We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of pr…

Computer Security

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

2026-05-31 · Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe, Hal Daumé arxiv

Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making reliable detection costly if it depends on …

Emergent Misalignment is Easy, Narrow Misalignment is Hard

2026-02-08 · Anna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel Nanda arxiv

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered surv…

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating

2026-06-08 · Sicheng Wang, Xiangyang Zhu, Han Wang, Zongrui Wang 외 arxiv

Prior work has shown that fine-tuning large language models on malicious or incorrect outputs in narrow domains can induce broad misalignment and harmful behavior, a phenomenon known as emergent misalignment. However, ef…

Characterizing the Consistency of the Emergent Misalignment Persona

2026-04-30 · Anietta Weckauff, Yuchen Zhang, Maksym Andriushchenko arxiv

Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignment (EM). While prior work has found a correlation between harmful be…