paper-with-me

Papers

Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment

2025-11-10 · Peng Zhang, Peijie Sun arxiv

Safety alignment instills in Large Language Models (LLMs) a critical capacity to refuse malicious requests. Prior works have modeled this refusal mechanism as a single linear direction in the activation space. We posit that this is an oversimplification that conflates two functionally distinct neural processes: the detection of harm and the execution of a refusal. In this work, we deconstruct this single representation into a Harm Detection Direction and a Refusal Execution Direction. Leveraging this fine-grained model, we introduce Differentiated Bi-Directional Intervention (DBDI), a new white-box framework that precisely neutralizes the safety alignment at critical layer. DBDI applies adaptive projection nullification to the refusal execution direction while suppressing the harm detection direction via direct steering. Extensive experiments demonstrate that DBDI outperforms prominent jailbreaking methods, achieving up to a 97.88\% attack success rate on models such as Llama-2. By providing a more granular and mechanistic framework, our work offers a new direction for the in-depth understanding of LLM safety alignment.

📄 PDF Abstract BibTeX arXiv:2511.06852

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Safety-Aware Role-Orchestrated Multi-Agent LLM Framework for Behavioral Health Communication Simulation

2026-03-31 · Ha Na Cho arxiv

Single-agent large language model (LLM) systems struggle to simultaneously support diverse conversational functions and maintain safety in behavioral health communication. We propose a safety-aware, role-orchestrated mul…

Safety-Critical Control under Multiple State and Input Constraints and Application to Fixed-Wing UAV

2023-08-08 · Donggeon David Oh, Dongjae Lee, H. Jin Kim

This study presents a framework to guarantee safety for a class of second-order nonlinear systems under multiple state and input constraints. To facilitate real-world applications, a safety-critical controller must consi…

Collision Avoidance

RepIt: Steering Language Models with Concept-Specific Refusal Vectors

2025-09-16 · Vincent Siu, Nathan W. Henry, Nicholas Crispino, Yang Liu 외 arxiv

Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework for isolating concept-specific representa…

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

2026-04-23 · Jason Tang, Stephen Law arxiv

Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized visual changes would plausibly shift human judgement for a specific scen…

Image Editing

LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models

2026-06-10 · Yan Hong, Kedong Xiu, Wei Li, Jun Lan 외 arxiv

We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving general capability under a compact interventi…