paper-with-me

홈 › Papers

SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution

2023-09-25 · Zhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng, Qinglong Wang, Zhan Qin, Zhibo Wang, Kui Ren

Advanced text-to-image models such as DALL$\cdot$E 2 and Midjourney possess the capacity to generate highly realistic images, raising significant concerns regarding the potential proliferation of unsafe content. This includes adult, violent, or deceptive imagery of political figures. Despite claims of rigorous safety mechanisms implemented in these models to restrict the generation of not-safe-for-work (NSFW) content, we successfully devise and exhibit the first prompt attacks on Midjourney, resulting in the production of abundant photorealistic NSFW images. We reveal the fundamental principles of such prompt attacks and suggest strategically substituting high-risk sections within a suspect prompt to evade closed-source safety measures. Our novel framework, SurrogatePrompt, systematically generates attack prompts, utilizing large language models, image-to-text, and image-to-image modules to automate attack prompt creation at scale. Evaluation results disclose an 88% success rate in bypassing Midjourney's proprietary safety filter with our attack prompts, leading to the generation of counterfeit images depicting political figures in violent scenarios. Both subjective and objective assessments validate that the images generated from our attack prompts present considerable safety hazards.

📄 PDF Abstract BibTeX arXiv:2309.14122

Code (0)

등록된 구현이 없습니다.

Tasks

Image to text

Similar Papers 제목 키워드 기반

Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

2025-09-21 · Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li 외 arxiv

Text-to-image (T2I) models have been widely applied in generating high-fidelity images across various domains. However, these models may also be abused to produce Not-Safe-for-Work (NSFW) content via jailbreak attacks. E…

Jailbreaks on Vision Language Model via Multimodal Reasoning

2026-01-29 · Aarush Noheria, Yuguang Yao arxiv

Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation. However, their outputs are highly sensitive to prompt variations, which can re…

Visual Question AnsweringText-to-Image GenerationMultimodal ReasoningImage Captioning

TokenProber: Jailbreaking Text-to-image Models via Fine-grained Word Impact Analysis

2025-05-11 · Longtian Wang, Xiaofei Xie, Tianlin Li, Yuhan Zhi 외

Text-to-image (T2I) models have significantly advanced in producing high-quality images. However, such models have the ability to generate images containing not-safe-for-work (NSFW) content, such as pornography, violence…

Sensitivity

Text is All You Need for Vision-Language Model Jailbreaking

2026-01-31 · Yihang Chen, Zhao Xu, Youyuan Jiang, Tianle Zheng 외 arxiv

Large Vision-Language Models (LVLMs) are increasingly equipped with robust safety safeguards to prevent responses to harmful or disallowed prompts. However, these defenses often focus on analyzing explicit textual inputs…

GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

2025-06-11 · Zilong Wang, Xiang Zheng, Xiaosen Wang, Bo wang 외

Text-to-image (T2I) models such as Stable Diffusion have advanced rapidly and are now widely used in content creation. However, these models can be misused to generate harmful content, including nudity or violence, posin…

Large Language ModelRed Teaming