paper-with-me

홈 › Papers

Enhancing Adversarial Attacks through Chain of Thought

2024-10-29 · Jingbo Su

Large language models (LLMs) have demonstrated impressive performance across various domains but remain susceptible to safety concerns. Prior research indicates that gradient-based adversarial attacks are particularly effective against aligned LLMs and the chain of thought (CoT) prompting can elicit desired answers through step-by-step reasoning. This paper proposes enhancing the robustness of adversarial attacks on aligned LLMs by integrating CoT prompts with the greedy coordinate gradient (GCG) technique. Using CoT triggers instead of affirmative targets stimulates the reasoning abilities of backend LLMs, thereby improving the transferability and universality of adversarial attacks. We conducted an ablation study comparing our CoT-GCG approach with Amazon Web Services auto-cot. Results revealed our approach outperformed both the baseline GCG attack and CoT prompting. Additionally, we used Llama Guard to evaluate potentially harmful interactions, providing a more objective risk assessment of entire conversations compared to matching outputs to rejection phrases. The code of this paper is available at https://github.com/sujingbo0217/CS222W24-LLM-Attack.

📄 PDF Abstract BibTeX arXiv:2410.21791

Code (1)

sujingbo0217/cs222w24-llm-attack 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models

2025-02-03 · Zhiyuan Xu, Joseph Gardiner, Sana Belguith

Large language models are typically trained on vast amounts of data during the pre-training phase, which may include some potentially harmful information. Fine-tuning attacks can exploit this by prompting the model to re…

Safety Alignment

Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection

2025-03-27 · Ryan Marinelli, Josef Pichlmeier, Tamas Bisztray

In this work, we propose a metric called Number of Thoughts (NofT) to determine the difficulty of tasks pre-prompting and support Large Language Models (LLMs) in production contexts. By setting thresholds based on the nu…

AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models

2025-09-29 · Zihao Zhu, Xinyu Wu, Gehan Hu, Siwei Lyu 외 arxiv

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in complex problem-solving through Chain-of-Thought (CoT) reasoning. However, the multi-step nature of CoT introduces new safety challenges that ext…

Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image

2024-02-22 · Zefeng Wang, Zhen Han, Shuo Chen, Fan Xue 외

Multimodal LLMs (MLLMs) with a great ability of text and image understanding have received great attention. To achieve better reasoning with MLLMs, Chain-of-Thought (CoT) reasoning has been widely explored, which further…

Adversarial RobustnessMultimodal ReasoningVisual Reasoning

Thought Purity: Defense Paradigm For Chain-of-Thought Attack

2025-07-16 · Zihao Xue, Zhen Bi, Long Ma, Zhenlin Hu 외

While reinforcement learning-trained Large Reasoning Models (LRMs, e.g., Deepseek-R1) demonstrate advanced reasoning capabilities in the evolving Large Language Models (LLMs) domain, their susceptibility to security thre…

reinforcement-learningReinforcement Learning