paper-with-me

Papers

DiffGuard: Text-Based Safety Checker for Diffusion Models

2024-11-25 · Massine El Khader, Elias Al Bouzidi, Abdellah Oumida, Mohammed Sbaihi, Eliott Binard, Jean-Philippe Poli, Wassila Ouerdane, Boussad ADDAD, Katarzyna Kapusta

Recent advances in Diffusion Models have enabled the generation of images from text, with powerful closed-source models like DALL-E and Midjourney leading the way. However, open-source alternatives, such as StabilityAI's Stable Diffusion, offer comparable capabilities. These open-source models, hosted on Hugging Face, come equipped with ethical filter protections designed to prevent the generation of explicit images. This paper reveals first their limitations and then presents a novel text-based safety filter that outperforms existing solutions. Our research is driven by the critical need to address the misuse of AI-generated content, especially in the context of information warfare. DiffGuard enhances filtering efficacy, achieving a performance that surpasses the best existing filters by over 14%.

📄 PDF Abstract BibTeX arXiv:2412.00064

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DiffGuard: Semantic Mismatch-Guided Out-of-Distribution Detection using Pre-trained Diffusion Models

2023-08-15 · ICCV 2023 1 · Ruiyuan Gao, Chenchen Zhao, Lanqing Hong, Qiang Xu

Given a classifier, the inherent property of semantic Out-of-Distribution (OOD) samples is that their contents differ from all legal classes in terms of semantics, namely semantic mismatch. There is a recent work that di…

Generative Adversarial NetworkOut-of-Distribution Detection

CROPS: Model-Agnostic Training-Free Framework for Safe Image Synthesis with Latent Diffusion Models

2025-01-09 · Junha Park, Ian Ryu, Jaehui Hwang, Hyungkeun Park 외

With advances in diffusion models, image generation has shown significant performance improvements. This raises concerns about the potential abuse of image generation, such as the creation of explicit or violent images, …

Image Generation

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

2025-07-30 · Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang 외 arxiv

Various (text) prompt filters and (image) safety checkers have been implemented to mitigate the misuse of Text-to-Image (T2I) models in creating Not-Safe-For-Work (NSFW) content. In order to expose potential security vul…

Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

2024-04-02 · Jiachen Ma, Yijiang Li, Zhiqing Xiao, Anda Cao 외

Text-to-image (T2I) models can be maliciously used to generate harmful content such as sexually explicit, unfaithful, and misleading or Not-Safe-for-Work (NSFW) images. Previous attacks largely depend on the availability…

Adversarial AttackImage GenerationText to Image GenerationText-to-Image Generation

MMA-Diffusion: MultiModal Attack on Diffusion Models

2023-11-29 · CVPR 2024 1 · Yijun Yang, Ruiyuan Gao, Xiaosen Wang, Tsung-Yi Ho 외

In recent years, Text-to-Image (T2I) models have seen remarkable advancements, gaining widespread adoption. However, this progress has inadvertently opened avenues for potential misuse, particularly in generating inappro…