paper-with-me

홈 › Papers

Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models

2025-05-03 · Gracjan Góral, Alicja Ziarko, Piotr Miłoś, Michał Nauman, Maciej Wołczyk, Michał Kosiński

We investigate the ability of Vision Language Models (VLMs) to perform visual perspective taking using a novel set of visual tasks inspired by established human tests. Our approach leverages carefully controlled scenes, in which a single humanoid minifigure is paired with a single object. By systematically varying spatial configurations - such as object position relative to the humanoid minifigure and the humanoid minifigure's orientation - and using both bird's-eye and surface-level views, we created 144 unique visual tasks. Each visual task is paired with a series of 7 diagnostic questions designed to assess three levels of visual cognition: scene understanding, spatial reasoning, and visual perspective taking. Our evaluation of several state-of-the-art models, including GPT-4-Turbo, GPT-4o, Llama-3.2-11B-Vision-Instruct, and variants of Claude Sonnet, reveals that while they excel in scene understanding, the performance declines significantly on spatial reasoning and further deteriorates on perspective-taking. Our analysis suggests a gap between surface-level object recognition and the deeper spatial and perspective reasoning required for complex visual tasks, pointing to the need for integrating explicit geometric representations and tailored training protocols in future VLM development.

📄 PDF Abstract BibTeX arXiv:2505.03821

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticObject RecognitionScene UnderstandingSpatial Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Visual Perspective Taking for Opponent Behavior Modeling

2021-05-11 · Boyuan Chen, Yuhang Hu, Robert Kwiatkowski, Shuran Song 외

In order to engage in complex social interaction, humans learn at a young age to infer what others see and cannot see from a different point-of-view, and learn to predict others' plans and behaviors. These abilities have…

Egocentric Bias in Vision-Language Models

2026-02-10 · Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao 외 arxiv

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vis…

Spatial Reasoning

You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction

2025-10-16 · Logan Lawrence, Oindrila Saha, Megan Wei, Chen Sun 외 arxiv

Despite the renewed interest in zero-shot visual classification due to the rise of Multimodal Large Language Models (MLLMs), the problem of evaluating free-form responses of auto-regressive models remains a persistent ch…

Fine-Grained Visual Recognition

An Interdisciplinary Perspective on Evaluation and Experimental Design for Visual Text Analytics: Position Paper

2022-09-23 · Kostiantyn Kucher, Nicole Sultanum, Angel Daza, Vasiliki Simaki 외

Appropriate evaluation and experimental design are fundamental for empirical sciences, particularly in data-driven fields. Due to the successes in computational modeling of languages, for instance, research outcomes are …

Experimental DesignPosition

Multi-View Active Fine-Grained Recognition

2022-06-02 · Ruoyi Du, Wenqing Yu, Heqing Wang, Dongliang Chang 외

As fine-grained visual classification (FGVC) being developed for decades, great works related have exposed a key direction -- finding discriminative local regions and revealing subtle differences. However, unlike identif…

Fine-Grained Image Classification