paper-with-me

홈 › Papers

CIVET: Systematic Evaluation of Understanding in VLMs

2025-06-05 · Massimo Rizzoli, Simone Alghisi, Olha Khomyn, Gabriel Roccabruna, Seyed Mahed Mousavi, Giuseppe Riccardi

While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding of VLMs, we study their capability regarding object properties and relations in a controlled and interpretable manner. To this scope, we introduce CIVET, a novel and extensible framework for systematiC evaluatIon Via controllEd sTimuli. CIVET addresses the lack of standardized systematic evaluation for assessing VLMs' understanding, enabling researchers to test hypotheses with statistical rigor. With CIVET, we evaluate five state-of-the-art VLMs on exhaustive sets of stimuli, free from annotation noise, dataset-specific biases, and uncontrolled scene complexity. Our findings reveal that 1) current VLMs can accurately recognize only a limited set of basic object properties; 2) their performance heavily depends on the position of the object in the scene; 3) they struggle to understand basic relations among objects. Furthermore, a comparative evaluation with human annotators reveals that VLMs still fall short of achieving human-level accuracy.

📄 PDF Abstract BibTeX arXiv:2506.05146

Code (0)

등록된 구현이 없습니다.

Tasks

Object

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Support is All You Need for Certified VAE Training

2025-04-16 · Changming Xu, Debangshu Banerjee, Deepak Vasisht, Gagandeep Singh

Variational Autoencoders (VAEs) have become increasingly popular and deployed in safety-critical applications. In such applications, we want to give certified probabilistic guarantees on performance under adversarial att…

All

NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

2025-04-04 · Kexin Tian, Jingrui Mao, Yunlong Zhang, Jiwan Jiang 외

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhib…

3d scene graph generationAutonomous DrivingGraph GenerationScene Graph Generation+1

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

2026-05-02 · Alejandro Aparcedo, Akash Kumar, Aaryan Garg, Dalton Pham 외 arxiv

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, m…

CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks

2024-06-20 · Jie Feng, Jun Zhang, Tianhui Liu, Xin Zhang 외

Recently, large language models (LLMs) with extensive general knowledge and powerful reasoning abilities have seen rapid development and widespread application. A systematic and reliable evaluation of LLMs or vision-lang…

General KnowledgeHuman DynamicsLanguage ModellingLarge Language Model

How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study

2026-04-08 · Roberto Brusnicki, Mattia Piccinini, Johannes Betz arxiv

Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poorly characterized, particularly regarding how input configurations affec…

Temporal SequencesAutonomous DrivingObject Detection