paper-with-me

홈 › Papers

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

2026-06-18 · Sihui Dai, Mann Patel arxiv

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harmful request, helpful response) with harmful compliance demonstrations (harmful request, helpful response) and testing three hypotheses about how demonstration composition drives harmful compliance. Across four models, we find that benign and harmful demonstrations are not interchangeable: benign demonstrations can either reduce or increase harmful compliance depending on the model. We further show that preference optimization is the critical training stage that prevents benign demonstrations from increasing harmful compliance, that demonstration ordering exhibits strong recency bias, and that models differ in how refusal interacts with in-context learning: some adopt demonstrated formatting even when refusing, while others override all in-context signals upon refusal. Taken together, this work moves beyond showing that demonstration-based jailbreaking works to characterizing how it works: what models extract from compliance demonstrations depends on demonstration content, ordering, and training methodology.

📄 PDF Abstract BibTeX arXiv:2606.20508

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

2023-10-05 · Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 외

Optimizing large language models (LLMs) for downstream use cases often involves the customization of pre-trained LLMs through further fine-tuning. Meta's open release of Llama models and OpenAI's APIs for fine-tuning GPT…

Red TeamingSafety Alignment

What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs

2025-05-26 · Sangyeop Kim, Yohan Lee, Yongwoo Song, Kimin Lee

We investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ). Our experiments utilize context length of up to 128K tokens. Through comprehensive analysis with various m…

Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights

2025-06-06 · Sooyung Choi, JaeHyeok Lee, Xiaoyuan Yi, Jing Yao 외

The application scope of Large Language Models (LLMs) continues to expand, leading to increasing interest in personalized LLMs that align with human values. However, aligning these models with individual values raises si…

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

2024-12-27 · Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay 외

Fine-tuning large language models (LLMs) for downstream tasks is a widely adopted approach, but it often leads to safety degradation in safety-aligned LLMs. Currently, many solutions address this issue by incorporating a…

Gradient Surgery for Safe LLM Fine-Tuning

2025-08-10 · Biao Yi, Jiahao Li, Baolei Zhang, Lihai Nie 외 arxiv

Fine-tuning-as-a-Service introduces a critical vulnerability where a few malicious examples mixed into the user's fine-tuning dataset can compromise the safety alignment of Large Language Models (LLMs). While a recognize…