paper-with-me

홈 › Papers

XBreaking: Explainable Artificial Intelligence for Jailbreaking LLMs

2025-04-30 · Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera, Vinod P

Large Language Models are fundamental actors in the modern IT landscape dominated by AI solutions. However, security threats associated with them might prevent their reliable adoption in critical application scenarios such as government organizations and medical institutions. For this reason, commercial LLMs typically undergo a sophisticated censoring mechanism to eliminate any harmful output they could possibly produce. In response to this, LLM Jailbreaking is a significant threat to such protections, and many previous approaches have already demonstrated its effectiveness across diverse domains. Existing jailbreak proposals mostly adopt a generate-and-test strategy to craft malicious input. To improve the comprehension of censoring mechanisms and design a targeted jailbreak attack, we propose an Explainable-AI solution that comparatively analyzes the behavior of censored and uncensored models to derive unique exploitable alignment patterns. Then, we propose XBreaking, a novel jailbreak attack that exploits these unique patterns to break the security constraints of LLMs by targeted noise injection. Our thorough experimental campaign returns important insights about the censoring mechanisms and demonstrates the effectiveness and performance of our attack.

📄 PDF Abstract BibTeX arXiv:2504.21700

Code (0)

등록된 구현이 없습니다.

Tasks

Explainable artificial intelligence

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models

2026-01-08 · Badhan Chandra Das, Md Tasnim Jawad, Joaquin Molto, M. Hadi Amini 외 arxiv

In recent years, the security vulnerabilities of Multi-modal Large Language Models (MLLMs) have become a serious concern in the Generative Artificial Intelligence (GenAI) research. These highly intelligent models, capabl…

JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models

2024-06-26 · Haibo Jin, Leyang Hu, Xinuo Li, Peiyan Zhang 외

The rapid evolution of artificial intelligence (AI) through developments in Large Language Models (LLMs) and Vision-Language Models (VLMs) has brought significant advancements across various technological domains. While …

LLM JailbreakSurvey

Explanation in Artificial Intelligence: Insights from the Social Sciences

2017-06-22 · Tim Miller

There has been a recent resurgence in the area of explainable artificial intelligence as researchers and practitioners seek to make their algorithms more understandable. Much of this research is focused on explicitly exp…

Explainable artificial intelligencePhilosophy

Towards Transparent AI: A Survey on Explainable Large Language Models

2025-06-26 · Avash Palikhe, Zhenyu Yu, Zichong Wang, Wenbin Zhang

Large Language Models (LLMs) have played a pivotal role in advancing Artificial Intelligence (AI). However, despite their achievements, LLMs often struggle to explain their decision-making processes, making them a 'black…

DecoderExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Survey

Deep Learning, Natural Language Processing, and Explainable Artificial Intelligence in the Biomedical Domain

2022-02-25 · Milad Moradi, Matthias Samwald

In this article, we first give an introduction to artificial intelligence and its applications in biology and medicine in Section 1. Deep learning methods are then described in Section 2. We narrow down the focus of the …

Explainable artificial intelligence