paper-with-me

Papers

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?

2024-12-04 · Shouwei Ruan, Hanqing Liu, Yao Huang, Xiaoqi Wang, Caixin Kang, Hang Su, Yinpeng Dong, Xingxing Wei

Vision Language Models (VLMs) have exhibited remarkable generalization capabilities, yet their robustness in dynamic real-world scenarios remains largely unexplored. To systematically evaluate VLMs' robustness to real-world 3D variations, we propose AdvDreamer, the first framework that generates physically reproducible adversarial 3D transformation (Adv-3DT) samples from single-view images. AdvDreamer integrates advanced generative techniques with two key innovations and aims to characterize the worst-case distributions of 3D variations from natural images. To ensure adversarial effectiveness and method generality, we introduce an Inverse Semantic Probability Objective that executes adversarial optimization on fundamental vision-text alignment spaces, which can be generalizable across different VLM architectures and downstream tasks. To mitigate the distribution discrepancy between generated and real-world samples while maintaining physical reproducibility, we design a Naturalness Reward Model that provides regularization feedback during adversarial optimization, preventing convergence towards hallucinated and unnatural elements. Leveraging AdvDreamer, we establish MM3DTBench, the first VQA dataset for benchmarking VLMs' 3D variations robustness. Extensive evaluations on representative VLMs with diverse architectures highlight that 3D variations in the real world may pose severe threats to model performance across various tasks.

📄 PDF Abstract BibTeX arXiv:2412.03002

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages

2020-04-28 · Katharina Kann, Ophélie Lacroix, Anders Søgaard

Part-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision - e.g., cross-lingual transfer, type-level supervision, or a combination thereof - have been report…

Cross-Lingual TransferPOSPOS Tagging

The Illusion-Illusion: Vision Language Models See Illusions Where There are None

2024-12-07 · Tomer Ullman

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something "really is" and how something "appears to be…

DiagnosticPhilosophy

Emergent Semantic Role Understanding in Language Models

2026-05-09 · Carla Griffiths, Mirco Musolesi arxiv

Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision they truly require. In particular, semantic role understanding ("wh…

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

2026-05-21 · Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou arxiv

Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motiv…

Visual Grounding

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports

2026-07-02 · Bingcong Yan, Chunlei Li, Jingliang Hu, Yilei Shi 외 arxiv

Large vision-language models (LVLMs) have achieved strong performance across many medical imaging tasks, yet their application to ultrasound remains limited due to its inherent complexity and variability. In this work, w…