paper-with-me

홈 › Papers

Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step

2024-10-04 · Wenxuan Wang, Kuiyi Gao, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Shuai Wang, Wenxiang Jiao, Zhaopeng Tu

Text-based image generation models, such as Stable Diffusion and DALL-E 3, hold significant potential in content creation and publishing workflows, making them the focus in recent years. Despite their remarkable capability to generate diverse and vivid images, considerable efforts are being made to prevent the generation of harmful content, such as abusive, violent, or pornographic material. To assess the safety of existing models, we introduce a novel jailbreaking method called Chain-of-Jailbreak (CoJ) attack, which compromises image generation models through a step-by-step editing process. Specifically, for malicious queries that cannot bypass the safeguards with a single prompt, we intentionally decompose the query into multiple sub-queries. The image generation models are then prompted to generate and iteratively edit images based on these sub-queries. To evaluate the effectiveness of our CoJ attack method, we constructed a comprehensive dataset, CoJ-Bench, encompassing nine safety scenarios, three types of editing operations, and three editing elements. Experiments on four widely-used image generation services provided by GPT-4V, GPT-4o, Gemini 1.5 and Gemini 1.5 Pro, demonstrate that our CoJ attack method can successfully bypass the safeguards of models for over 60% cases, which significantly outperforms other jailbreaking methods (i.e., 14%). Further, to enhance these models' safety against our CoJ attack method, we also propose an effective prompting-based method, Think Twice Prompting, that can successfully defend over 95% of CoJ attack. We release our dataset and code to facilitate the AI safety research.

📄 PDF Abstract BibTeX arXiv:2410.03869

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

2026-05-08 · Jianming Chen, Yawen Wang, Junjie Wang, Zhe Liu 외 arxiv

Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, this capability introduces a new safety attack surface: harmful outputs m…

Poisoned LangChain: Jailbreak LLMs by LangChain

2024-06-26 · Ziqiu Wang, Jun Liu, Shengkai Zhang, Yang Yang

With the development of natural language processing (NLP), large language models (LLMs) are becoming increasingly popular. LLMs are integrating more into everyday life, raising public concerns about their security vulner…

RAGRetrievalRetrieval-augmented Generation

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

2026-02-10 · Jiacheng Hou, Yining Sun, Ruochong Jin, Haochen Han 외 arxiv

Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-…

Multimodal ReasoningImage Editing

Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models

2025-04-15 · Jiangtao Liu, Zhaoxin Wang, Handing Wang, Cong Tian 외

Recent advancements in Text-to-Image (T2I) generation have significantly enhanced the realism and creativity of generated images. However, such powerful generative capabilities pose risks related to the production of ina…

DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing

2025-02-17 · Yi Wang, Fenghua Weng, Sibei Yang, Zhan Qin 외

Large Language Models (LLMs) are widely applied in decision making, but their deployment is threatened by jailbreak attacks, where adversarial users manipulate model behavior to bypass safety measures. Existing defense m…

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model+3