paper-with-me

Papers

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

2025-07-01 · Tao Lin, Gen Li, Yilei Zhong, Yanwen Zou, Yuxin Du, Jiting Liu, Encheng Gu, Bo Zhao arxiv

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models (VLMs), which excel at semantic understanding due to large-scale image and text pretraining. However, existing VLMs typically lack precise spatial understanding capabilities, as they are primarily tuned on 2D image-text pairs without 3D supervision. To address this limitation, recent approaches have incorporated explicit 3D inputs such as point clouds or depth maps, but this necessitates additional depth sensors or pre-trained depth estimation models, which may yield defective results. In contrast, our work introduces a plug-and-play module that implicitly incorporates 3D geometry features into VLA models by leveraging an off-the-shelf visual geometry foundation model. This integration provides the model with depth-aware visual representations, improving its ability to understand the geometric structure of the scene and the spatial relationships among objects from RGB images alone. We evaluate our method on a set of spatially challenging tasks in both simulation and the real world. Extensive evaluations show that our method significantly improves the performance of state-of-the-art VLA models across diverse scenarios.

📄 PDF Abstract BibTeX arXiv:2507.00416

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationPoint Clouds

Similar Papers 제목 키워드 기반

Grounded 3D-Aware Spatial Vision-Language Modeling

2026-05-28 · An-Chieh Cheng, Yang Fu, Yatai Ji, Ligeng Zhu 외 arxiv

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introdu…

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

2025-09-06 · Ruixun Liu, Lingyu Kong, Derun Li, Hang Zhao arxiv

Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key …

Multimodal ReasoningTrajectory PlanningAutonomous Driving

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

2026-05-14 · Tao Lin, Yuxin Du, Jiting Liu, Nuobei Zhu 외 arxiv

Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise s…

Point Clouds

Weakly Supervised Relative Spatial Reasoning for Visual Question Answering

2021-09-04 · ICCV 2021 10 · Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral

Vision-and-language (V\&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. O…

Question AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)+1

Exploring Spatial Schema Intuitions in Large Language and Vision Models

2024-02-01 · Philipp Wicke, Lennart Wachowiak

Despite the ubiquity of large language models (LLMs) in AI research, the question of embodiment in LLMs remains underexplored, distinguishing them from embodied systems in robotics where sensory perception directly infor…

Language ModelingLanguage Modelling