paper-with-me

Papers

HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate

2025-12-09 · Shenzhe Zhu arxiv

Large language models (LLMs) are equipped with safety mechanisms to detect and block harmful queries, yet current alignment approaches primarily focus on overtly dangerous content and overlook more subtle threats. However, users can often disguise harmful intent through covert rephrasing that preserves malicious objectives while appearing benign, which creates a significant gap in existing safety training data. To address this limitation, we introduce HarmTransform, a multi-agent debate framework for systematically transforming harmful queries into stealthier forms while preserving their underlying harmful intent. Our framework leverages iterative critique and refinement among multiple agents to generate high-quality, covert harmful query transformations that can be used to improve future LLM safety alignment. Experiments demonstrate that HarmTransform significantly outperforms standard baselines in producing effective query transformations. At the same time, our analysis reveals that debate acts as a double-edged sword: while it can sharpen transformations and improve stealth, it may also introduce topic shifts and unnecessary complexity. These insights highlight both the promise and the limitations of multi-agent debate for generating comprehensive safety training data.

📄 PDF Abstract BibTeX arXiv:2512.23717

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

2026-01-08 · Huawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji 외 arxiv

Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risks. Domain-specific datasets of harmful prompts remain scarce and still l…

Text is All You Need for Vision-Language Model Jailbreaking

2026-01-31 · Yihang Chen, Zhao Xu, Youyuan Jiang, Tianle Zheng 외 arxiv

Large Vision-Language Models (LVLMs) are increasingly equipped with robust safety safeguards to prevent responses to harmful or disallowed prompts. However, these defenses often focus on analyzing explicit textual inputs…

Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment

2026-03-12 · Zhiyu Xue, Zimo Qi, Guangliang Liu, Bocheng Chen 외 arxiv

Safety alignment aims to ensure that large language models (LLMs) refuse harmful requests by post-training on harmful queries paired with refusal answers. Although safety alignment is widely adopted in industry, the over…

Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding

2025-09-13 · Seongho Joo, Hyukhun Koh, Kyomin Jung arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their potential misuse for harmful purposes remains a significant concern. To strengthen defenses against such vulnerabilit…

Alignment faking in large language models

2024-12-18 · Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger 외

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Cla…

Large Language Model