paper-with-me

Papers

B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

2024-03-14 · Hao Zhang, Wenqi Shao, Hong Liu, Yongqiang Ma, Ping Luo, Yu Qiao, Nanning Zheng, Kaipeng Zhang

Large Vision-Language Models (LVLMs) have shown significant progress in responding well to visual-instructions from users. However, these instructions, encompassing images and text, are susceptible to both intentional and inadvertent attacks. Despite the critical importance of LVLMs' robustness against such threats, current research in this area remains limited. To bridge this gap, we introduce B-AVIBench, a framework designed to analyze the robustness of LVLMs when facing various Black-box Adversarial Visual-Instructions (B-AVIs), including four types of image-based B-AVIs, ten types of text-based B-AVIs, and nine types of content bias B-AVIs (such as gender, violence, cultural, and racial biases, among others). We generate 316K B-AVIs encompassing five categories of multimodal capabilities (ten tasks) and content bias. We then conduct a comprehensive evaluation involving 14 open-source LVLMs to assess their performance. B-AVIBench also serves as a convenient tool for practitioners to evaluate the robustness of LVLMs against B-AVIs. Our findings and extensive experimental results shed light on the vulnerabilities of LVLMs, and highlight that inherent biases exist even in advanced closed-source LVLMs like GeminiProVision and GPT-4V. This underscores the importance of enhancing the robustness, security, and fairness of LVLMs. The source code and benchmark are available at https://github.com/zhanghao5201/B-AVIBench.

📄 PDF Abstract BibTeX arXiv:2403.09346

Code (1)

zhanghao5201/b-avibench 공식 구현 jax

Tasks

FairnessLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

V$^2$R-Bench: Holistically Evaluating LVLM Robustness to Fundamental Visual Variations

2025-04-23 · Zhiyuan Fan, Yumeng Wang, Sandeep Polisetty, Yi R. Fung

Large Vision Language Models (LVLMs) excel in various vision-language tasks. Yet, their robustness to visual variations in position, scale, orientation, and context that objects in natural scenes inevitably exhibit due t…

Dataset GenerationObject RecognitionPosition

On Evaluating Adversarial Robustness of Large Vision-Language Models

2023-05-26 · NeurIPS 2023 11 · Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang 외

Large vision-language models (VLMs) such as GPT-4 have achieved unprecedented performance in response generation, especially with visual inputs, enabling more creative and adaptable interaction than large language models…

Adversarial Robustnessmultimodal generationResponse Generation

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

2024-09-20 · Hao Cheng, Erjia Xiao, Chengyuan Yu, Zhao Yao 외

Recently, driven by advancements in Multimodal Large Language Models (MLLMs), Vision Language Action Models (VLAMs) are being proposed to achieve better performance in open-vocabulary scenarios for robotic manipulation t…

Vision-Language-Action

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

2025-11-18 · Huiyi Chen, Jiawei Peng, Dehai Min, Changchang Sun 외 arxiv

Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-world applications. However, existing robustness benchmarks typically focu…

LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation

2024-10-07 · Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang 외

Building on the advancements of Large Language Models (LLMs) and Vision Language Models (VLMs), recent research has introduced Vision-Language-Action (VLA) models as an integrated solution for robotic manipulation tasks.…

Vision-Language-Action