paper-with-me

홈 › Papers

Measuring How (Not Just Whether) VLMs Build Common Ground

2025-09-04 · Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani arxiv

Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gradually develop shared understanding through ongoing communication. We introduce a four-metric suite (grounding efficiency, content alignment, lexical adaptation, and human-likeness) to systematically evaluate VLM performance in interactive grounding contexts. We deploy the suite on 150 self-play sessions of interactive referential games between three proprietary VLMs and compare them with human dyads. All three models diverge from human patterns on at least three metrics, while GPT4o-mini is the closest overall. We find that (i) task success scores do not indicate successful grounding and (ii) high image-utterance alignment does not necessarily predict task success. Our metric suite and findings offer a framework for future research on VLM grounding.

📄 PDF Abstract BibTeX arXiv:2509.03805

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Understanding Figurative Meaning through Explainable Visual Entailment

2024-05-02 · Arkadiy Saakyan, Shreyas Kulkarni, Tuhin Chakrabarty, Smaranda Muresan

Large Vision-Language Models (VLMs) have demonstrated strong capabilities in tasks requiring a fine-grained understanding of literal meaning in images and text, such as visual question-answering or visual entailment. How…

Question AnsweringVisual EntailmentVisual Question Answering

Do Image Editing Models Understand Lighting?

2026-06-25 · Tim Küchler, Johann-Friedrich Feiden, Matthias Nießner, Carsten Rother arxiv

While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems possess an intrinsic knowledge of real-world lighting. Existing bench…

Image Editing

UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models

2024-07-25 · Xinyu Pi, Mingyuan Wu, Jize Jiang, Haozhen Zheng 외

Smaller-scale Vision-Langauge Models (VLMs) often claim to perform on par with larger models in general-domain visual grounding and question-answering benchmarks while offering advantages in computational efficiency and …

Computational EfficiencyQuestion AnsweringVisual Grounding

Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images

2025-05-12 · Elisei Rykov, Kseniia Petrushina, Kseniia Titova, Anton Razzhigaev 외

Measuring how real images look is a complex task in artificial intelligence research. For example, an image of a boy with a vacuum cleaner in a desert violates common sense. We introduce a novel method, which we call Thr…

Common Sense Reasoning

Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights

2025-06-19 · Yuan Zhong, Ruinan Jin, Xiaoxiao Li, Qi Dou

Medical vision-language models (VLMs) leverage large-scale pretraining for diverse imaging tasks but require substantial computational and data resources. Meanwhile, common or general-purpose VLMs (e.g., CLIP, LLaVA), th…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)