paper-with-me

홈 › Papers

LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models

2024-07-23 · Shi Lin, Hongming Yang, Rongchang Li, Xun Wang, Changting Lin, Wenpeng Xing, Meng Han

The rapid development of Large Language Models (LLMs) has brought impressive advancements across various tasks. However, despite these achievements, LLMs still pose inherent safety risks, especially in the context of jailbreak attacks. Most existing jailbreak methods follow an input-level manipulation paradigm to bypass safety mechanisms. Yet, as alignment techniques improve, such attacks are becoming increasingly detectable. In this work, we identify an underexplored threat vector: the model's internal reasoning process, which can be manipulated to elicit harmful outputs in a more stealthy way. To explore this overlooked attack surface, we propose a novel black-box jailbreak attack method, Analyzing-based Jailbreak (ABJ). ABJ comprises two independent attack paths: textual and visual reasoning attacks, which exploit the model's multimodal reasoning capabilities to bypass safety mechanisms, comprehensively exposing vulnerabilities in its reasoning chain. We conduct extensive experiments on ABJ across various open-source and closed-source LLMs, VLMs, and RLMs. In particular, ABJ achieves high attack success rate (ASR) (82.1% on GPT-4o-2024-11-20) with exceptional attack efficiency (AE) among all target models, showcasing its remarkable attack effectiveness, transferability, and efficiency. Our work reveals a new type of safety risk and highlights the urgent need to mitigate implicit vulnerabilities in the model's reasoning process.

📄 PDF Abstract BibTeX arXiv:2407.16205

Code (1)

theshi-1128/ABJ-Attack 공식 구현

Tasks

Multimodal ReasoningPrompt EngineeringVisual Reasoning

Methods 이 논문이 사용한 방법론

AE An autoencoder is a type of artificial neural network used to learn efficient data codings in an unsupervised manner. The aim of an autoencoder is to learn a representation…

Similar Papers 제목 키워드 기반

Multi-round jailbreak attack on large language models

2024-10-15 · Yihua Zhou, Xiaochuan Shi

Ensuring the safety and alignment of large language models (LLMs) with human values is crucial for generating responses that are beneficial to humanity. While LLMs have the capability to identify and avoid harmful querie…

Prompt Engineering

Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks

2024-08-29 · Tom Gibbs, Ethan Kosak-Hine, George Ingebretsen, Jason Zhang 외

Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In th…

Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

2023-12-07 · Yanrui Du, Sendong Zhao, Ming Ma, Yuhan Chen 외

Extensive work has been devoted to improving the safety mechanism of Large Language Models (LLMs). However, LLMs still tend to generate harmful responses when faced with malicious instructions, a phenomenon referred to a…

Dark LLMs: The Growing Threat of Unaligned AI Models

2025-05-15 · Michael Fire, Yitzhak Elbazis, Adi Wasenstein, Lior Rokach

Large Language Models (LLMs) rapidly reshape modern life, advancing fields from healthcare to education and beyond. However, alongside their remarkable capabilities lies a significant threat: the susceptibility of these …

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

2025-08-22 · Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin 외 arxiv

With the development of Large Language Models (LLMs), numerous efforts have revealed their vulnerabilities to jailbreak attacks. Although these studies have driven the progress in LLMs' safety alignment, it remains uncle…