paper-with-me

홈 › Papers

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

2026-07-01 · Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah arxiv

Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g., FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with structured evaluation, this illusion breaks: on TIFA (Text-to-Image Faithfulness evaluation with Question Answering), safety-aligned models suffer substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships. To diagnose the source of this gap, we analyze the text-encoder prompt embedding space and uncover semantic collapse, a contraction of embedding spread coupled with distortion of inter-prompt similarity structure, which strongly correlates with structured utility loss. Guided by this insight, we propose StructureAware Geometric Regularization (SAGE), a safety alignment objective that explicitly preserves embedding spread and inter-prompt relational structure during adaptation. Our method restores structured utility (TIFA +5.0% over prior state-of-the-art) while maintaining strong safety performance and competitive coarse-grained utility scores. Our source code and trained models are available at https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/.

📄 PDF Abstract BibTeX arXiv:2607.00402

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Illusions in Humans and AI: How Visual Perception Aligns and Diverges

2025-08-17 · Jianyi Yang, Junyi Ye, Ankan Dash, Guiling Wang arxiv

By comparing biological and artificial perception through the lens of illusions, we highlight critical differences in how each system constructs visual reality. Understanding these divergences can inform the development …

SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions

2026-03-24 · Jinzhe Tu, Ruilei Guo, Zihan Guo, Junxiao Yang 외 arxiv

Recent works have shown that multimodal large language models (MLLMs) are highly vulnerable to hidden-pattern visual illusions, where the hidden content is imperceptible to models but obvious to humans. This deficiency h…

Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning

2026-02-14 · Yanbo Wang, Minzheng Wang, Jian Liang, Lu Wang 외 arxiv

While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures. For safety alignment, the core challenge lies in the inherent trade-off b…

Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

2024-05-22 · Weixiang Zhao, Yulin Hu, Zhuojun Li, Yang Deng 외

Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generat…

Safety Alignment

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings

2025-11-26 · Fatemeh Akbarian, Anahita Baninajjar, Yingyi Zhang, Ananth Balashankar 외 arxiv

Multi-modal foundation models align images, text, and other modalities in a shared embedding space but remain vulnerable to adversarial illusions [35], where imperceptible perturbations disrupt cross-modal alignment and …