paper-with-me

Papers

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

2025-10-13 · Nikita Afonin, Nikita Andriianov, Vahagn Hovhannisyan, Nikhil Bageshpura, Kyle Liu, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Oleg Rogov, Elena Tutubalina, Alexander Panchenko, Mikhail Seleznyov arxiv

Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning, these findings were limited to finetuning and activation steering, leaving out in-context learning (ICL). We therefore ask: does EM emerge in ICL? We find that it does: across four model families (Gemini, Kimi-K2, Grok, and Qwen), narrow in-context examples cause models to produce misaligned responses to benign, unrelated queries. With 16 in-context examples, EM rates range from 1% to 24% depending on model and domain, appearing with as few as 2 examples. Neither larger model scale nor explicit reasoning provides reliable protection, and larger models are typically even more susceptible. Next, we formulate and test a hypothesis, which explains in-context EM as conflict between safety objectives and context-following behavior. Consistent with this, instructing models to prioritize safety reduces EM while prioritizing context-following increases it. These findings establish ICL as a previously underappreciated vector for emergent misalignment that resists simple scaling-based solutions.

📄 PDF Abstract BibTeX arXiv:2510.11288

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantic Containment as a Fundamental Property of Emergent Misalignment

2026-02-02 · Rohan Saxena arxiv

Fine-tuning language models on narrowly harmful data causes emergent misalignment (EM) -- behavioral failures extending far beyond training distributions. Recent work demonstrates compartmentalization of misalignment beh…

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

2026-05-12 · Baris Askin, Muhammed Ustaomeroglu, Anupam Nayak, Gauri Joshi 외 arxiv

Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better under…

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

2025-07-04 · Jeremiah Giordani arxiv

Recent work has shown that fine-tuning large language models (LLMs) on code with security vulnerabilities can result in misaligned and unsafe behaviors across broad domains. These results prompted concerns about the emer…

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

2026-04-28 · Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan 외 arxiv

Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the tra…

Reinforcement Learning Can Amplify Emergent Misalignment from Harmless Rewards

2026-05-29 · Magnus Jørgenvåg, David Kaczér, Lasse Ruttert, Marvin Gülhan 외 arxiv

Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While EM has been extensively studied in the supervised fine-tuning…

Reinforcement Learning