ProSPer: Probing Human and Neural Network Language Model Understanding of Spatial Perspective
Understanding perspectival language is important for applications like dialogue systems and human-robot interaction. We propose a probe task that explores how well language models understand spatial perspective. We present a dataset for evaluating perspective inference in English, ProSPer, and use it to explore how humans and Transformer-based language models infer perspective. Although the best bidirectional model performs similarly to humans, they display different strengths: humans outperform neural networks in conversational contexts, while RoBERTa excels at written genres.
Code (1)
Tasks
Language ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language mod…
Spatial ReasoningCorruption and Wealth: Unveiling a national prosperity syndrome in Europe
Data mining revealed a cluster of economic, psychological, social and cultural indicators that in combination predicted corruption and wealth of European nations. This prosperity syndrome of self-reliant citizens, effici…
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation …
Self-Supervised LearningScene UnderstandingVisuospatial Perspective Taking in Multimodal Language Models
As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or s…
Scene UnderstandingIs BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language Understanding
Most humans use visual imagination to understand and reason about language, but models such as BERT reason about language using knowledge acquired during text-only pretraining. In this work, we investigate whether vision…
Knowledge ProbingLanguage ModellingNatural Language UnderstandingVisual Reasoning