paper-with-me

홈 › Papers

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

2025-10-15 · Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Zhuoguang Chen, Tao Jiang, Hang Zhao arxiv

Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requiring precise spatial reasoning due to limited spatial reasoning inherited from Vision-Language Models (VLMs). Existing VLAs rely on extensive action-data pretraining to ground VLMs in 3D space, which reduces training efficiency and is still insufficient for accurate spatial understanding. In this work, we present DepthVLA, a simple yet effective VLA architecture that explicitly incorporates spatial awareness through a pretrained depth prediction module. DepthVLA adopts a mixture-of-transformers design that unifies a VLM, a depth transformer, and an action expert with fully shared attentions, forming an end-to-end model with enhanced spatial reasoning. Extensive evaluations in both real-world and simulated environments show that DepthVLA outperforms state-of-the-art approaches, achieving 78.5% vs. 65.0% progress in real-world tasks, 94.9% vs. 93.6% in the LIBERO simulator, and 74.8% vs. 58.8% in the Simpler simulator. Our code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2510.13375

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

GEM: Generative Supervision Helps Embodied Intelligence

2026-05-27 · Ruowen Zhao, Bangguo Li, Zuyan Liu, Yinan Liang 외 arxiv

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-l…

Composition Vision-Language Understanding via Segment and Depth Anything Model

2024-06-07 · Mingxiao Huo, Pengliang Ji, Haotian Lin, Junchen Liu 외

We introduce a pioneering unified library that leverages depth anything, segment anything models to augment neural comprehension in language-vision model zero-shot understanding. This library synergizes the capabilities …

Question AnsweringVisual Question Answering (VQA)

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

2025-10-16 · Yixuan Li, Yuhui Chen, Mingcai Zhou, Haoran Li 외 arxiv

Spatial perception and reasoning are crucial for Vision-Language-Action (VLA) models to accomplish fine-grained manipulation tasks. However, existing approaches often lack the ability to understand and reason over the es…

Spatial Reasoning

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

2025-12-26 · Shengliang Deng, Mi Yan, Yixin Zheng, Jiayi Su 외 arxiv

While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoint robustness. This limitation largely stems from the reliance on pretrai…

Spatial ReasoningDepth Estimation

Enhancing Vision-Based Policies with Omni-View and Cross-Modality Knowledge Distillation for Mobile Robots

2026-03-21 · Kai Li, Shiyu Zhao arxiv

Vision-based policies are widely applied in robotics for tasks such as manipulation and locomotion. On lightweight mobile robots, however, they face a trilemma of limited scene transferability, restricted onboard computa…

Knowledge Distillation