paper-with-me

Papers

Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks

2025-11-13 · Xuancun Lu, Jiaxiang Chen, Shilin Xiao, Zizhi Jin, Zhangrui Chen, Hanwen Yu, Bohan Qian, Ruochen Zhou, Xiaoyu Ji, Wenyuan Xu arxiv

Vision-Language-Action (VLA) models revolutionize robotic systems by enabling end-to-end perception-to-action pipelines that integrate multiple sensory modalities, such as visual signals processed by cameras and auditory signals captured by microphones. This multi-modality integration allows VLA models to interpret complex, real-world environments using diverse sensor data streams. Given the fact that VLA-based systems heavily rely on the sensory input, the security of VLA models against physical-world sensor attacks remains critically underexplored. To address this gap, we present the first systematic study of physical sensor attacks against VLAs, quantifying the influence of sensor attacks and investigating the defenses for VLA models. We introduce a novel "Real-Sim-Real" framework that automatically simulates physics-based sensor attack vectors, including six attacks targeting cameras and two targeting microphones, and validates them on real robotic systems. Through large-scale evaluations across various VLA architectures and tasks under varying attack parameters, we demonstrate significant vulnerabilities, with susceptibility patterns that reveal critical dependencies on task types and model designs. We further develop an adversarial-training-based defense that enhances VLA robustness against out-of-distribution physical perturbations caused by sensor attacks while preserving model performance. Our findings expose an urgent need for standardized robustness benchmarks and mitigation strategies to secure VLA deployments in safety-critical environments.

📄 PDF Abstract BibTeX arXiv:2511.10008

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models

2024-08-02 · Simone Caldarella, Massimiliano Mancini, Elisa Ricci, Rahaf Aljundi

Vision-Language Models (VLMs) combine visual and textual understanding, rendering them well-suited for diverse tasks like generating image captions and answering visual questions across various domains. However, these ca…

Image Captioning

Phantom: A Unified Face-Swap Deepfake Protection Framework with Latent and Spatial Constraints

2026-06-30 · Jungkon Kim, Cheolseung Jung, Jong-Min Choi, Juseong Lee arxiv

Face-swapping deepfakes pose an escalating threat to personal privacy by enabling unauthorized identity manipulation. While adversarial approaches have demonstrated success against black-box face recognition (FR) models,…

Face Recognition

Learning opening books in partially observable games: using random seeds in Phantom Go

2016-07-08 · Tristan Cazenave, Jialin Liu, Fabien Teytaud, Olivier Teytaud

Many artificial intelligences (AIs) are randomized. One can be lucky or unlucky with the random seed; we quantify this effect and show that, maybe contrarily to intuition, this is far from being negligible. Then, we appl…

A quantum active learning algorithm for sampling against adversarial attacks

2019-12-06 · P. A. M. Casares, M. A. Martin-Delgado

Adversarial attacks represent a serious menace for learning algorithms and may compromise the security of future autonomous systems. A theorem by Khoury and Hadfield-Menell (KH), provides sufficient conditions to guarant…

Active Learning

microPhantom: Playing microRTS under uncertainty and chaos

2020-05-22 · Florian Richoux

This competition paper presents microPhantom, a bot playing microRTS and participating in the 2020 microRTS AI competition. microPhantom is based on our previous bot POAdaptive which won the partially observable track of…

Decision MakingDecision Making Under Uncertainty