paper-with-me

Papers

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks

2026-04-14 · Yingying Zhao, Chengyin Hu, Qike Zhang, Xin Li, Xin Wang, Yiwei Wei, Jiujiang Guo, Jiahuan Long, Tingsong Jiang, Wen Yao arxiv

Vision-Language Models (VLMs) have shown remarkable performance, yet their security remains insufficiently understood. Existing adversarial studies focus almost exclusively on the digital setting, leaving physical-world threats largely unexplored. As VLMs are increasingly deployed in real environments, this gap becomes critical, since adversarial perturbations must be physically realizable. Despite this practical relevance, physical attacks against VLMs have not been systematically studied. Such attacks may induce recognition failures and further disrupt multimodal reasoning, leading to severe semantic misinterpretation in downstream tasks. Therefore, investigating physical attacks on VLMs is essential for assessing their real-world security risks. To address this gap, we propose Multimodal Semantic Lighting Attacks (MSLA), the first physically deployable adversarial attack framework against VLMs. MSLA uses controllable adversarial lighting to disrupt multimodal semantic understanding in real scenes, attacking semantic alignment rather than only task-specific outputs. Consequently, it degrades zero-shot classification performance of mainstream CLIP variants while inducing severe semantic hallucinations in advanced VLMs such as LLaVA and BLIP across image captioning and visual question answering (VQA). Extensive experiments in both digital and physical domains demonstrate that MSLA is effective, transferable, and practically realizable. Our findings provide the first evidence that VLMs are highly vulnerable to physically deployable semantic attacks, exposing a previously overlooked robustness gap and underscoring the urgent need for physical-world robustness evaluation of VLMs.

📄 PDF Abstract BibTeX arXiv:2604.12833

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringMultimodal ReasoningAdversarial AttackImage Captioning

Similar Papers 제목 키워드 기반

SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters

2025-09-18 · Abdarahmane Traore, Éric Hervet, Andy Couturier arxiv

Recent advances in vision-language models (VLMs) have enabled powerful multimodal reasoning, but state-of-the-art approaches typically rely on extremely large models with prohibitive computational and memory requirements…

Multimodal ReasoningSpatial Reasoning

GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

2026-07-08 · Peter Bohm, Saimunur Rahman, Abdelwahed Khamis, Sagun Man Singh Shrestha 외 arxiv

Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask…

Visual NavigationRobot Navigation

Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction

2026-09-07 · Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang 외 arxiv

Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable vis…

Zero-shot Generalization

OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

2025-11-03 · Heyu Guo, Shanmu Wang, Ruichun Ma, Shiqi Jiang 외 arxiv

Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models rely solely on RGB cameras, limiting their…

Towards deployment-centric multimodal AI beyond vision and language

2025-04-04 · Xianyuan Liu, Jiayang Zhang, Shuo Zhou, Thijs L. van der Plas 외

Multimodal artificial intelligence (AI) integrates diverse types of data via machine learning to improve understanding, prediction, and decision-making across disciplines such as healthcare, science, and engineering. How…

Decision Making