paper-with-me

홈 › Papers

Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models

2025-08-12 · Wei Cai, Jian Zhao, Yuchu Jiang, Tianle Zhang, Xuelong Li arxiv

Large Vision-Language Models face growing safety challenges with multimodal inputs. This paper introduces the concept of Implicit Reasoning Safety, a vulnerability in LVLMs. Benign combined inputs trigger unsafe LVLM outputs due to flawed or hidden reasoning. To showcase this, we developed Safe Semantics, Unsafe Interpretations, the first dataset for this critical issue. Our demonstrations show that even simple In-Context Learning with SSUI significantly mitigates these implicit multimodal threats, underscoring the urgent need to improve cross-modal implicit reasoning.

📄 PDF Abstract BibTeX arXiv:2508.08926

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interpretation modeling: Social grounding of sentences by reasoning over their implicit moral judgments

2023-11-27 · Liesbeth Allein, Maria Mihaela Truşcǎ, Marie-Francine Moens

The social and implicit nature of human communication ramifies readers' understandings of written sentences. Single gold-standard interpretations rarely exist, challenging conventional assumptions in natural language pro…

Sentence

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

2026-07-07 · Yuanmin Huang, Zhenfei Zhang, Mi Zhang, Geng Hong 외 arxiv

Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, …

SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers

2026-04-02 · Xiang Yang, Feifei Li, Mi Zhang, Geng Hong 외 arxiv

Recent Text-to-Image (T2I) models based on rectified-flow transformers (e.g., SD3, FLUX) achieve high generative fidelity but remain vulnerable to unsafe semantics, especially when triggered by multi-token interactions. …

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

2026-06-05 · Xiang Yang, Feifei Li, Mi Zhang, Geng Hong 외 arxiv

Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful content remains a critical challenge, particu…

Image Generation

DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation

2026-03-23 · Binhong Tan, Zhaoxin Wang, Handing Wang arxiv

Text-to-Image (T2I) diffusion models have demonstrated strong generation ability, but their potential to generate unsafe content raises significant safety concerns. Existing inference-time defense methods typically perfo…

Text-to-Image Generation