paper-with-me

홈 › Papers

Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification

2025-08-28 · Xiangtao Meng, Yingkai Dong, Ning Yu, Li Wang, Zheng Li, Shanqing Guo arxiv

Text-to-image (T2I) generative models have achieved remarkable visual fidelity, yet remain vulnerable to generating unsafe content. Existing safety defenses typically intervene internally within the generative model, but suffer from severe concept entanglement, leading to degradation of benign generation quality, a trade-off we term the Safety Tax. To overcome this limitation, we advocate a paradigm shift from destructive internal editing to external safety rectification. Following this principle, we propose SafePatch, a structurally isolated safety module that performs external, interpretable rectification without modifying the base model. The core backbone of SafePatch is architecturally instantiated as a trainable clone of the base model's encoder, allowing it to inherit rich semantic priors and maintain representation consistency. To enable interpretable safety rectification, we construct a strictly aligned counterfactual safety dataset (ACS) for differential supervision training. Across nudity and multi-category benchmarks and recent adversarial prompt attacks, SafePatch achieves robust unsafe suppression (7% unsafe on I2P) while preserving image quality and semantic alignment.

📄 PDF Abstract BibTeX arXiv:2508.21099

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

2023-05-23 · Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes 외

State-of-the-art Text-to-Image models like Stable Diffusion and DALLE$\cdot$2 are revolutionizing how people generate visual content. At the same time, society has serious concerns about how adversaries can exploit such …

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

2026-06-03 · Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers 외 arxiv

Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture. In this…

Text-to-Video Generation

ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection

2025-11-24 · Ruize Ma, Minghong Cai, Yilei Jiang, Jiaming Han 외 arxiv

Recent progress in video generative models has enabled the creation of high-quality videos from multimodal prompts that combine text and images. While these systems offer enhanced controllability, they also introduce new…

Video Generation

Mitigating Covertly Unsafe Text within Natural Language Systems

2022-10-17 · Alex Mei, Anisha Kabir, Sharon Levy, Melanie Subbiah 외

An increasingly prevalent problem for intelligent technologies is text safety, as uncontrolled systems may generate recommendations to their users that lead to injury or life-threatening consequences. However, the degree…

ProGuard: Towards Proactive Multimodal Safeguard

2025-12-29 · Shaohan Yu, Lijun Li, Chenyang Si, Lu Sheng 외 arxiv

The rapid evolution of generative models has led to a continuous emergence of multimodal safety risks, exposing the limitations of existing defense methods. To address these challenges, we propose ProGuard, a vision-lang…

Reinforcement Learning