paper-with-me

홈 › Papers

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

2026-05-26 · Yifan Jiang, Ruoxi Ning, Sheng Yao, Freda Shi arxiv

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental image context in lexical judgments. We use human concreteness and imagery ratings because they span words with varying expected visual relevance, from abstract and low-imagery words to concrete and high-imagery words. We find that real-image contexts do not yield consistent gains and often hurt alignment with human ratings, most sharply when visual evidence is least relevant. Through probing and canonical correlation analysis, complemented by an attribution case study, we find that real-image contexts are associated with representational shifts and greater sensitivity to spurious visual cues, coinciding with weaker recoverability of the targeted lexical properties. We further show that instructing models to focus solely on textual content at inference time can reduce this degradation, with the clearest gains on these vulnerable subsets. Our findings suggest that current instruction-tuned VLMs need better calibration of when visual context should inform lexical judgments.

📄 PDF Abstract BibTeX arXiv:2605.27315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Divergences in Color Perception between Deep Neural Networks and Humans

2023-09-11 · Ethan O. Nadler, Elise Darragh-Ford, Bhargav Srinivasa Desikan, Christian Conaway 외

Deep neural networks (DNNs) are increasingly proposed as models of human vision, bolstered by their impressive performance on image classification and object recognition tasks. Yet, the extent to which DNNs capture funda…

image-classificationImage ClassificationImage SegmentationObject Recognition+2

ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs

2026-04-07 · Zhipin Wang, Christoph Leiter, Christian Frey, Mohamed Hesham Ibrahim Abdalla 외 arxiv

Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cultural values in language models are almost entirely text-only, leaving …

Large Language Models as 'Hidden Persuaders': Fake Product Reviews are Indistinguishable to Humans and Machines

2025-06-16 · Weiyao Meng, John Harvey, James Goulding, Chris James Carter 외

Reading and evaluating product reviews is central to how most people decide what to buy and consume online. However, the recent emergence of Large Language Models and Generative Artificial Intelligence now means writing …

DiffuSyn Bench: Evaluating Vision-Language Models on Real-World Complexities with Diffusion-Generated Synthetic Benchmarks

2024-06-06 · Haokun Zhou, Yipeng Hong

This study assesses the ability of Large Vision-Language Models (LVLMs) to differentiate between AI-generated and human-generated images. It introduces a new automated benchmark construction method for this evaluation. T…

Image GenerationRetrievalScript Generation

Non-identifiability of Explanations from Model Behavior in Deep Networks of Image Authenticity Judgments

2026-04-08 · Icaro Re Depaolini, Uri Hasson arxiv

Deep neural networks can predict human judgments, but this does not imply that they rely on human-like information or reveal the cues underlying those judgments. Prior work has addressed this issue using attribution heat…