paper-with-me

홈 › Papers

Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

2023-11-15 · Haoran Wang, Kai Shu

To ensure AI safety, instruction-tuned Large Language Models (LLMs) are specifically trained to ensure alignment, which refers to making models behave in accordance with human intentions. While these models have demonstrated commendable results on various safety benchmarks, the vulnerability of their safety alignment has not been extensively studied. This is particularly troubling given the potential harm that LLMs can inflict. Existing attack methods on LLMs often rely on poisoned training data or the injection of malicious prompts. These approaches compromise the stealthiness and generalizability of the attacks, making them susceptible to detection. Additionally, these models often demand substantial computational resources for implementation, making them less practical for real-world applications. In this work, we study a different attack scenario, called Trojan Activation Attack (TA^2), which injects trojan steering vectors into the activation layers of LLMs. These malicious steering vectors can be triggered at inference time to steer the models toward attacker-desired behaviors by manipulating their activations. Our experiment results on four primary alignment tasks show that TA^2 is highly effective and adds little or no overhead to attack efficiency. Additionally, we discuss potential countermeasures against such activation attacks.

📄 PDF Abstract BibTeX arXiv:2311.09433

Code (1)

wang2226/backdoor-activation-attack 공식 구현 pytorch

Tasks

Red TeamingSafety Alignment

Similar Papers 제목 키워드 기반

Instance-Level Trojan Attacks on Visual Question Answering via Adversarial Learning in Neuron Activation Space

2023-04-02 · Yuwei Sun, Hideya Ochiai, Jun Sakuma

Trojan attacks embed perturbations in input data leading to malicious behavior in neural network models. A combination of various Trojans in different modalities enables an adversary to mount a sophisticated attack on mu…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Causality Analysis for Evaluating the Security of Large Language Models

2023-12-13 · Wei Zhao, Zhe Li, Jun Sun

Large Language Models (LLMs) such as GPT and Llama2 are increasingly adopted in many safety-critical applications. Their security is thus essential. Even with considerable efforts spent on reinforcement learning from hum…

Red Teaming

Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models

2026-02-19 · Manuel Wirth arxiv

As Large Language Models (LLMs) are increasingly integrated into automated decision-making pipelines, specifically within Human Resources (HR), the security implications of Indirect Prompt Injection (IPI) become critical…

Dormant Neural Trojans

2022-11-02 · Feisi Fu, Panagiota Kiourti, Wenchao Li

We present a novel methodology for neural network backdoor attacks. Unlike existing training-time attacks where the Trojaned network would respond to the Trojan trigger after training, our approach inserts a Trojan that …

TeleLoRA: Teleporting Model-Specific Alignment Across LLMs

2025-03-26 · Xiao Lin, Manoj Acharya, Anirban Roy, Susmit Jha

Mitigating Trojans in Large Language Models (LLMs) is one of many tasks where alignment data is LLM specific, as different LLMs have different Trojan triggers and trigger behaviors to be removed. In this paper, we introd…

model