paper-with-me

홈 › Papers

Egocentric Bias in Vision-Language Models

2026-02-10 · Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo arxiv

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models. The task requires simulating 180-degree rotations of 2D character strings from another agent's perspective, isolating spatial transformation from 3D scene complexity. Evaluating 103 VLMs reveals systematic egocentric bias: the vast majority perform below chance, with roughly three-quarters of errors reproducing the camera viewpoint. Control experiments expose a compositional deficit--models achieve high theory-of-mind accuracy and above-chance mental rotation in isolation, yet fail catastrophically when integration is required. This dissociation indicates that current VLMs lack the mechanisms needed to bind social awareness to spatial operations, suggesting fundamental limitations in model-based spatial reasoning. FlipSet provides a cognitively grounded testbed for diagnosing perspective-taking capabilities in multimodal systems.

📄 PDF Abstract BibTeX arXiv:2602.15892

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

EgoNCE++: Do Egocentric Video-Language Models Really Understand Hand-Object Interactions?

2024-05-28 · Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song 외

Egocentric video-language pretraining is a crucial paradigm to advance the learning of egocentric hand-object interactions (EgoHOI). Despite the great success on existing testbeds, these benchmarks focus more on closed-s…

Action RecognitionAttributeIn-Context LearningMulti-Instance Retrieval

PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence

2025-12-18 · Xiaopeng Lin, Shijie Lian, Bin Yu, Ruoqi Yang 외 arxiv

Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception and action. Vision Language Models (VLMs…

Domain Adaptive Egocentric Person Re-identification

2021-03-08 · Ankit Choudhary, Deepak Mishra, Arnab Karmakar

Person re-identification (re-ID) in first-person (egocentric) vision is a fairly new and unexplored problem. With the increase of wearable video recording devices, egocentric data becomes readily available, and person re…

Person Re-IdentificationStyle Transfer

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

2025-12-30 · Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 외 arxiv

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for u…

Referring Video Object Segmentation

Spatial-Conditioned Reasoning in Long-Egocentric Videos

2026-01-26 · James Tribble, Hao Wang, Si-En Hong, Chaoyi Zhou 외 arxiv

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and…

Spatial ReasoningVisual Navigation