paper-with-me

홈 › Papers

Visuospatial Perspective Taking in Multimodal Language Models

2026-03-04 · Jonathan Prunty, Seraphina Zhang, Patrick Quinn, Jianxun Lian, Xing Xie, Lucy Cheke arxiv

As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or static scene understanding, leaving visuospatial perspective-taking (VPT) underexplored. We adapt two evaluation tasks from human studies: the Director Task, assessing VPT in a referential communication paradigm, and the Rotating Figure Task, probing perspective-taking across angular disparities. Across tasks, MLMs show pronounced deficits in Level 2 VPT, which requires inhibiting one's own perspective to adopt another's. These results expose critical limitations in current MLMs' ability to represent and reason about alternative perspectives, with implications for their use in collaborative contexts.

📄 PDF Abstract BibTeX arXiv:2603.23510

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

Gender differences in the neural structures of (non) emotional perspective-taking: A tDCS study

2022-05-24 · Shahab Mahdian, Michael A. Nitsche, Vahid Nejati

The perspective-taking in social cognition is an ability that makes third-person judgments about the intentions, beliefs and thoughts of others. We aimed to investigate gender differences in the neural structure differen…

Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts

2025-05-18 · Qi Feng

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models …

Spatial Reasoning

Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs

2026-04-17 · Rohit Sinha, Aditya Kanade, Sai Srinivas Kancheti, Vineeth N Balasubramanian 외 arxiv

Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's E…

Egocentric Bias in Vision-Language Models

2026-02-10 · Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao 외 arxiv

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vis…

Spatial Reasoning

Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models

2026-01-23 · Bridget Leonard, Scott O. Murray arxiv

Multimodal language models (MLMs) perform well on semantic vision-language tasks but fail at spatial reasoning that requires adopting another agent's visual perspective. These errors reflect a persistent egocentric bias …

Spatial Reasoning