paper-with-me

홈 › Papers

VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety

2025-10-21 · Shruti Palaskar, Leon Gatys, Mona Abdelrahman, Mar Jacobo, Larry Lindsey, Rutika Moharir, Gunnar Lund, Yang Xu, Navid Shiee, Jeffrey Bigham, Charles Maalouf, Joseph Yitan Cheng arxiv

Safety evaluation of multimodal foundation models often treats vision and language inputs separately, missing risks from joint interpretation where benign content becomes harmful in combination. Existing approaches also fail to distinguish clearly unsafe content from borderline cases, leading to problematic over-blocking or under-refusal of genuinely harmful content. We present Vision Language Safety Understanding (VLSU), a comprehensive framework to systematically evaluate multimodal safety through fine-grained severity classification and combinatorial analysis across 17 distinct safety patterns. Using a multi-stage pipeline with real-world images and human annotation, we construct a large-scale benchmark of 8,187 samples spanning 15 harm categories. Our evaluation of eleven state-of-the-art models reveals systematic joint understanding failures: while models achieve 90%-plus accuracy on clear unimodal safety signals, performance degrades substantially to 20-55% when joint image-text reasoning is required to determine the safety label. Most critically, 34% of errors in joint image-text safety classification occur despite correct classification of the individual modalities, further demonstrating absent compositional reasoning capabilities. Additionally, we find that models struggle to balance refusing unsafe content while still responding to borderline cases that deserve engagement. For example, we find that instruction framing can reduce the over-blocking rate on borderline content from 62.4% to 10.4% in Gemini-1.5, but only at the cost of under-refusing on unsafe content with refusal rate dropping from 90.8% to 53.9%. Overall, our framework exposes weaknesses in joint image-text understanding and alignment gaps in current models, and provides a critical test bed to enable the next milestones in research on robust vision-language safety.

📄 PDF Abstract BibTeX arXiv:2510.18214

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

2026-08-06 · Hong Jiang, Junnan Zhu, Jingwang Huang, Xiao Sun 외 arxiv

Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings…

Reinforcement Learning

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

2025-11-21 · Chi Zhang, Jiepeng Wang, Youming Wang, Yuanzhi Liang 외 arxiv

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: th…

Text-to-Image Generation

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

2024-07-16 · Zehan Wang, Ziang Zhang, Hang Zhang, Luping Liu 외

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipeli…

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation

2026-08-13 · Yunhao Bai, Zhongwei Qiu, Guangyu Guo, Yiming Huang 외 arxiv

Clinical intelligence requires estimating a patient's underlying condition from incomplete observations rather than learning isolated mappings from scans to answers. Volumetric medical images provide dense observations o…

AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities

2024-12-18 · CVPR 2025 1 · Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu

Geospatial models must adapt to the diversity of Earth observation data in terms of resolutions, scales, and modalities. However, existing approaches expect fixed input configurations, which limits their practical applic…

Change DetectionDiversityEarth Observation