paper-with-me

Papers

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

2026-05-21 · Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou arxiv

Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising observation that removing a substantial fraction of image tokens only degrades model performance very slightly on a widely used hallucination benchmark, we systematically investigate this mismatch in a set of open-source VLMs. Our analysis spans multiple levels of granularity, spanning global visual degradation, localized occlusion, question reformulation, answer-space expansion, and decision-level analyses beyond standard accuracy. We further complement these behavioral results with a layer-wise analysis of vision-token geometry. Throughout the experiments, we find that although VLMs do incorporate visual input, their predictions are less sensitive to the loss of fine-grained visual evidence that standard accuracy should have suggested. Even when the final prediction remains unchanged, the model's internal support for the correct answer may already be weakened. We further complement a representation-level analysis, which shows increasing similarity among visual tokens in deeper layers, providing a possible explanation for our findings. Together, these results suggest that current benchmarks are not sufficient to reliably evaluate fine-grained visual grounding in VLMs.

📄 PDF Abstract BibTeX arXiv:2605.22903

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Vision: looking and seeing through our brain's information bottleneck

2025-03-24 · Li Zhaoping

Our brain recognizes only a tiny fraction of sensory input, due to an information processing bottleneck. This blinds us to most visual inputs. Since we are blind to this blindness, only a recent framework highlights this…

Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models

2025-12-17 · Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 외 arxiv

Multimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in connecting vision and language, yet their proficiency in fundamental visual reasoning tasks remains limited. This limitation …

Multimodal ReasoningVisual Reasoning

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

2025-11-15 · Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin 외 arxiv

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work…

SeeingSounds: Learning Audio-to-Visual Alignment via Text

2025-10-10 · Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo 외 arxiv

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on …

Image Generation

Strategic Collusion of LLM Agents: Market Division in Multi-Commodity Competitions

2024-09-19 · Ryan Y. Lin, Siddhartha Ojha, Kevin Cai, Maxwell F. Chen

Machine-learning technologies are seeing increased deployment in real-world market scenarios. In this work, we explore the strategic behaviors of large language models (LLMs) when deployed as autonomous agents in multi-c…