paper-with-me

홈 › Papers

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

2026-04-20 · Haiweng Xu, Sipeng Zheng, Hao Luo, Wanpeng Zhang, Ziheng Xi, Zongqing Lu arxiv

Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between standard benchmark success and true embodied reasoning, raising the question of whether these high scores reflect genuine cognitive capability. To address this gap, we introduce BeTTER, a diagnostic Benchmark for Testing True Embodied Reasoning in robotic policies. BeTTER applies targeted causal interventions (e.g., spatial layout shifts, temporal extrapolation) while enforcing kinematic isolation to explicitly decouple high-level reasoning failures from low-level execution limits. Through systematic evaluation, we reveal that state-of-the-art VLAs catastrophically fail in dynamic scenarios, exhibiting severe lexical-kinematic shortcuts, behavioral inertia, and semantic feature collapse. Crucially, our mechanistic analysis traces these symptoms to fundamental architectural bottlenecks - such as capacity compression and myopic downsampling - which systematically degrade the model's foundational semantic representation. We demonstrate that highly static evaluation protocols effectively mask this degradation by allowing optimization to overfit to sensorimotor priors. Supported by real-world robotic validation, our findings confirm that this representational breakdown is not a simulation artifact, highlighting the critical need for future VLA paradigms to resolve the structural tension between high-frequency control and high-level reasoning.

📄 PDF Abstract BibTeX arXiv:2604.18000

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

2026-07-30 · Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang 외 arxiv

Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as math…

IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

2024-03-23 · HAZ Sameen Shahgir, Khondker Salman Sayeed, Abhik Bhattacharjee, Wasi Uddin Ahmad 외

The advent of Vision Language Models (VLM) has allowed researchers to investigate the visual understanding of a neural network using natural language. Beyond object classification and detection, VLMs are capable of visua…

Common Sense ReasoningIn-Context LearningMultiple-choiceObject Localization+2

Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions

2026-03-31 · Xuesong Wang, Harry Wang arxiv

Vision-language models (VLMs) exhibit a systematic bias when confronted with classic optical illusions: they overwhelmingly predict the illusion as "real" regardless of whether the image has been counterfactually modifie…

Image ManipulationSpatial ReasoningImage Compression

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

2026-02-02 · Wenjin Hou, Wei Liu, Han Hu, Xiaoxiao Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. However, these evaluations typically rely on s…

Visual Reasoning

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

2026-05-18 · Yajing Zhou, Xiangyu Kong arxiv

While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains hampered by a "Cartesian Illusion" - a reliance on text-based probabil…

Spatial Reasoning