paper-with-me

홈 › Papers

JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization

2026-03-22 · Haolun Zheng, Yu He, Tailun Chen, Shuo Shao, Zhixuan Chu, Hongbin Zhou, Lan Tao, Zhan Qin, Kui Ren arxiv

Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive.

📄 PDF Abstract BibTeX arXiv:2603.21208

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

2025-01-29 · Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan 외

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model s…

Image GenerationInstruction FollowingText to Image Generation+2

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

2026-05-18 · Jun-Yu Pan, Yansen Wang, Enze Zhang, Bao-Liang Lu 외 arxiv

Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce, leading existing methods to align neural…

Visual Grounding

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

2025-06-22 · Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen 외

Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabili…

GPUImage GenerationLanguage ModelingLanguage Modelling+4

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

2026-06-29 · Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun 외 hf

Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding …

Zero-shot Generalization

A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice

2025-12-23 · Yaowei Bai, Ruiheng Zhang, Yu Lei, Xuhua Duan 외 arxiv

A global shortage of radiologists has been exacerbated by the significant volume of chest X-ray workloads, particularly in primary care. Although multimodal large language models show promise, existing evaluations predom…