paper-with-me

홈 › Papers

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks

2026-04-06 · Jia Chengyu, AprilPyone MaungMaung, Huy H. Nguyen, Jinyin Chen, Isao Echizen arxiv

Recent advances in vision-language models (VLMs) trained on web-scale image-text pairs have enabled impressive zero-shot transfer across a diverse range of visual tasks. However, comprehensive and independent evaluation beyond standard benchmarks is essential to understand their robustness, limitations, and real-world applicability. This paper presents a systematic evaluation framework for VLMs under natural adversarial scenarios for diverse downstream tasks, which has been overlooked in previous evaluation works. We evaluate a wide range of VLMs (CLIP, robust CLIP, BLIP2, and SigLIP2) on curated adversarial datasets (typographic attacks, ImageNet-A, and natural language-induced adversarial examples). We measure the natural adversarial performance of selected VLMs for zero-shot image classification, semantic segmentation, and visual question answering. Our analysis reveals that robust CLIP models can amplify natural adversarial vulnerabilities, and CLIP models significantly reduce performance for natural language-induced adversarial examples. Additionally, we provide interpretable analyses to identify failure modes. We hope our findings inspire future research in robust and fair multimodal pattern recognition.

📄 PDF Abstract BibTeX arXiv:2604.04473

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-Shot Image ClassificationVisual Question AnsweringSemantic Segmentation

Similar Papers 제목 키워드 기반

Pathological Truth Bias in Vision-Language Models

2025-09-14 · Yash Thube arxiv

Vision Language Models (VLMs) are improving quickly, but standard benchmarks can hide systematic failures that reduce real world trust. We introduce MATS (Multimodal Audit for Truthful Spatialization), a compact behavior…

Law and the Emerging Political Economy of Algorithmic Audits

2024-04-03 · Petros Terzis, Michael Veale, Noëlle Gaumann

For almost a decade now, scholarship in and beyond the ACM FAccT community has been focusing on novel and innovative ways and methodologies to audit the functioning of algorithmic systems. Over the years, this research i…

Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs

2026-02-12 · Longyuan Zhu, Hairan Hua, Linlin Miao, Bing Zhao arxiv

Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting have eroded the authority of standard ben…

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

2026-04-23 · Tasnim Kabir, Dmytro Kurdydyk, Aadi Palnitkar, Liam Dorn 외 arxiv

Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed through shortcut strategies, short-duration cues, lexical priors, da…

Question Answering

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

2026-04-27 · Xinming Tu, Tianze Wang, Yingzhou, Lu 외 arxiv

As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts t…