paper-with-me

홈 › Papers

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

2026-09-14 · Aashiq Muhamed, Mona T. Diab, Virginia Smith hf

Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defending against these attacks typically requires computationally expensive safety finetuning for every new checkpoint. We introduce Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning. Our approach is based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction. Rather than trying to hide the true refusal circuitry, DDO actively injects a high-magnitude, nonlinear decoy signal into the network's MLP neurons. When an attacker attempts to locate the refusal direction, the decoy corrupts their estimator, tricking them into ablating a harmless orthogonal feature while the actual safety mechanism remains intact. We prove a spectral bound formalizing this effect and evaluate DDO across six model families, achieving <10% ASR under standard RFA. On Llama-3-8B-Instruct, DDO remains comparable to trained defenses under adaptive multi-phase attacks (65% vs. 58% worst-case ASR) and reduces Heretic weight-level attack ASR from 88.7% to 18%, all at 30 to 450 times lower optimization cost per configuration than the trained baselines.

📄 PDF Abstract BibTeX arXiv:2609.16204

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

2026-08-02 · Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy 외 arxiv

We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack…

DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity

2026-01-18 · Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi, Yash Shah 외 arxiv

Multimodal Large Language Models (MLLMs) can directly consume exam documents, threatening conventional assessments and academic integrity. We present DoPE (Decoy-Oriented Perturbation Encapsulation), a document-layer def…

Towards Causal Models for Adversary Distractions

2021-04-21 · Ron Alford, Andy Applebaum

Automated adversary emulation is becoming an indispensable tool of network security operators in testing and evaluating their cyber defenses. At the same time, it has exposed how quickly adversaries can propagate through…

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

2026-07-29 · Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao 외 arxiv

Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observati…

Adversarial Decoys: Misdirecting Attention-Based Defenses in ViT

2026-07-08 · Giulia Marchiori Pietrosanti, Giulio Rossolini, Giorgio Buttazzo arxiv

Vision Transformers (ViTs) remain vulnerable to localized adversarial attacks, e.g., adversarial patches, while recent test-time defenses mitigate them by suppressing image tokens with abnormally high attention scores. T…