paper-with-me

홈 › Papers

Failures in Perspective-taking of Multimodal AI Systems

2024-09-20 · Bridget Leonard, Kristin Woodard, Scott O. Murray

This study extends previous research on spatial representations in multimodal AI systems. Although current models demonstrate a rich understanding of spatial information from images, this information is rooted in propositional representations, which differ from the analog representations employed in human and animal spatial cognition. To further explore these limitations, we apply techniques from cognitive and developmental science to assess the perspective-taking abilities of GPT-4o. Our analysis enables a comparison between the cognitive development of the human brain and that of multimodal AI, offering guidance for future research and model development.

📄 PDF Abstract BibTeX arXiv:2409.13929

Code (1)

bridgetleonard2/PerspectiveTaking 공식 구현

Similar Papers 제목 키워드 기반

Visuospatial Perspective Taking in Multimodal Language Models

2026-03-04 · Jonathan Prunty, Seraphina Zhang, Patrick Quinn, Jianxun Lian 외 arxiv

As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or s…

Scene Understanding

PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision

2025-11-11 · Sabrina Patania, Luca Annese, Anita Pellegrini, Silvia Serino 외 arxiv

Recent advances in Large Language Models (LLMs) and multimodal foundation models have significantly broadened their application in robotics and collaborative systems. However, effective multi-agent interaction necessitat…

3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction

2023-10-23 · Mehdi Fatan, Emanuele Mincato, Dimitra Pintzou, Mariella Dimiccoli

Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have sho…

Egocentric Bias in Vision-Language Models

2026-02-10 · Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao 외 arxiv

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vis…

Spatial Reasoning

Mass-Producing Failures of Multimodal Systems with Language Models

2023-06-21 · NeurIPS 2023 11 · Shengbang Tong, Erik Jones, Jacob Steinhardt

Deployed multimodal systems can fail in ways that evaluators did not anticipate. In order to find these failures before deployment, we introduce MultiMon, a system that automatically identifies systematic failures -- gen…

Language ModelingLanguage ModellingSelf-Driving Cars