paper-with-me

홈 › Papers

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

2026-06-22 · Yifei Zhao, Qian Lou, Mengxin Zheng arxiv

Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, where perception or reasoning errors can lead to unsafe decisions or actions. Although many red-teaming methods have been developed to probe VLM vulnerabilities, their evaluation remains fragmented across datasets, metrics, and threat models, making direct comparison difficult and obscuring whether observed differences arise from stronger attacks, more vulnerable models, or incompatible evaluation settings. Existing chatbot-centric red-teaming benchmarks mainly standardize jailbreak and content-safety evaluation, but they do not systematically capture physically grounded functional failures or cover red-teaming methods that target physical-world VLMs. This raises the key challenge of comparing diverse attack methods under a unified protocol while targeting the same scenario-specific failures. We introduce REALM, to our knowledge the first unified red-teaming benchmark for physical-world VLMs. REALM integrates 12 red-teaming methods, 3 model-agnostic defenses, and 13 VLMs under a practical black-box threat model with shared datasets and metrics. To align adversarial objectives across attack families, REALM introduces an agentic target-generation pipeline that constructs shared, scenario-specific, and physically grounded attack objectives for each scene, enabling fair comparison of diverse red-teaming methods under aligned adversarial goals. Our evaluation shows that text and typographic injection attacks induce the most failures, multimodal co-optimization yields the strongest visual-perturbation transfer, single-pass attacks approach iterative methods at much lower cost, and model scale alone does not confer adversarial robustness. Code is available at https://github.com/UCF-ML-Research/REALM.

📄 PDF Abstract BibTeX arXiv:2606.23892

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

2025-06-18 · Yining Hong, Rui Sun, Bingxuan Li, Xingcheng Yao 외

AI agents today are mostly siloed - they either retrieve and reason over vast amount of digital information and knowledge obtained online; or interact with the physical world through embodied perception, planning and act…

Physical Backdoor: Towards Temperature-based Backdoor Attacks in the Physical World

2024-04-30 · CVPR 2024 1 · Wen Yin, Jian Lou, Pan Zhou, Yulai Xie 외

Backdoor attacks have been well-studied in visible light object detection (VLOD) in recent years. However, VLOD can not effectively work in dark and temperature-sensitive scenarios. Instead, thermal infrared object detec…

Objectobject-detectionObject Detection

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming

2025-06-04 · Xiang Zheng, Xingjun Ma, Wei-Bin Lee, Cong Wang

Red teaming has proven to be an effective method for identifying and mitigating vulnerabilities in Large Language Models (LLMs). Reinforcement Fine-Tuning (RFT) has emerged as a promising strategy among existing red team…

Red Teaming

Towards Red Teaming in Multimodal and Multilingual Translation

2024-01-29 · Christophe Ropers, David Dale, Prangthip Hansanti, Gabriel Mejia Gonzalez 외

Assessing performance in Natural Language Processing is becoming increasingly complex. One particular challenge is the potential for evaluation datasets to overlap with training data, either directly or indirectly, which…

Machine TranslationRed TeamingTranslation

AT-Drone: Benchmarking Adaptive Teaming in Multi-Drone Pursuit

2025-02-13 · Yang Li, Junfan Chen, Feng Xue, Jiabin Qiu 외

Adaptive teaming-the capability of agents to effectively collaborate with unfamiliar teammates without prior coordination-is widely explored in virtual video games but overlooked in real-world multi-robot contexts. Yet, …

BenchmarkingEdge-computing