paper-with-me

홈 › Papers

VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?

2026-03-09 · Minkyu Kim, Sangheon Lee, Dongmin Park arxiv

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced reasoning required for real-world applications. In this work, we introduce VLM-SubtleBench, a benchmark designed to evaluate VLMs on subtle comparative reasoning. Our benchmark covers ten difference types - Attribute, State, Emotion, Temporal, Spatial, Existence, Quantity, Quality, Viewpoint, and Action - and curate paired question-image sets reflecting these fine-grained variations. Unlike prior benchmarks restricted to natural image datasets, our benchmark spans diverse domains, including industrial, aerial, and medical imagery. Through extensive evaluation of both proprietary and open-source VLMs, we reveal systematic gaps between model and human performance across difference types and domains, and provide controlled analyses highlighting where VLMs' reasoning sharply deteriorates. Together, our benchmark and findings establish a foundation for advancing VLMs toward human-level comparative reasoning.

📄 PDF Abstract BibTeX arXiv:2603.07888

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detection

Similar Papers 제목 키워드 기반

MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?

2024-10-06 · Guanzhen Li, Yuxi Xie, Min-Yen Kan

Humans perform visual perception at multiple levels, including low-level object recognition and high-level semantic interpretation such as behavior understanding. Subtle differences in low-level details can lead to subst…

Object Recognition

Enhancing Visual Classification using Comparative Descriptors

2024-11-08 · Hankyeol Lee, Gawon Seo, Wonseok Choi, Geunyoung Jung 외

The performance of vision-language models (VLMs), such as CLIP, in visual classification tasks, has been enhanced by leveraging semantic knowledge from large language models (LLMs), including GPT. Recent studies have sho…

Classificationzero-shot-classificationZero-Shot Learning

VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models

2025-05-28 · Chahat Raj, Bowen Wei, Aylin Caliskan, Antonios Anastasopoulos 외

While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less attention. Existing VLM bias studies often focus on portrait-style images and…

Decision MakingQuestion AnsweringVisual Question Answering (VQA)

CIVET: Systematic Evaluation of Understanding in VLMs

2025-06-05 · Massimo Rizzoli, Simone Alghisi, Olha Khomyn, Gabriel Roccabruna 외

While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding …

Object

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding

2026-05-19 · Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Hung-Ting Su 외 arxiv

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real-world applications requiring nuanced in…

Video Question AnsweringSpatial Reasoning