paper-with-me

홈 › Papers

Vision-language models lag human performance on physical dynamics and intent reasoning

2026-01-04 · Tianjun Gu, Jingyu Gong, Zhizhong Zhang, Yuan Xie, Lizhuang Ma, Xin Tan, Athanasios V arxiv

Spatial intelligence is central to embodied cognition, yet contemporary AI systems still struggle to reason about physical interactions in open-world human environments. Despite strong performance on controlled benchmarks, vision-language models often fail to jointly model physical dynamics, reference frames, and the latent human intentions that drive spatial change. We introduce Teleo-Spatial Intelligence (TSI), a reasoning capability that links spatiotemporal change to goal-directed structure. To evaluate TSI, we present EscherVerse, a large-scale open-world resource built from 11,328 real-world videos, including an 8,000-example benchmark and a 35,963-example instruction-tuning set. Across 27 state-of-the-art vision-language models and an independent analysis of first-pass human responses from 11 annotators, we identify a persistent teleo-spatial reasoning gap: the strongest proprietary model achieves 57.26\% overall accuracy, far below first-pass human performance, which ranges from 84.81\% to 95.14\% with a mean of 90.62\%. Fine-tuning on real-world, intent-aware data narrows this gap for open-weight models, but does not close it. EscherVerse provides a diagnostic testbed for purpose-aware spatial reasoning and highlights a critical gap between pattern recognition and human-level understanding in embodied AI.

📄 PDF Abstract BibTeX arXiv:2601.01547

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

PhysBrain 1.0 Technical Report

2026-05-14 · Shijie Lian, Bin Yu, Xiaopeng Lin, Changti Wu 외 arxiv

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale hu…

RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation

2026-02-11 · Yuhao Chen, Zhihao Zhan, Xiaoxin Lin, Zijian Song 외 arxiv

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings. This mismatch creates a substantial real…

Spatial Reasoning

Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models

2026-04-25 · Jimin Lee, Huiwon Jang, Myungkyu Koo, Jungwoo Park 외 arxiv

Humans understand and interact with the real world by relying on diverse physical feedback beyond visual perception. Motivated by this, recent approaches attempt to incorporate physical sensory signals into Vision-Langua…

Physion: Evaluating Physical Prediction from Vision in Humans and Machines

2021-06-15 · Daniel M. Bear, Elias Wang, Damian Mrowca, Felix J. Binder 외

While current vision algorithms excel at many challenging tasks, it is unclear how well they understand the physical dynamics of real-world environments. Here we introduce Physion, a dataset and benchmark for rigorously …

Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation

2026-03-02 · Jisoo Kim, Jungbin Cho, Sanghyeok Chu, Ananya Bal 외 arxiv

Humans learn not only how their bodies move, but also how the surrounding world responds to their actions. In contrast, while recent Vision-Language-Action (VLA) models exhibit impressive semantic understanding, they oft…