paper-with-me

홈 › Papers

3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks

2025-05-09 · Vineet Bhat, Yu-Hsiang Lan, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami

Robotic manipulation in 3D requires learning an $N$ degree-of-freedom joint space trajectory of a robot manipulator. Robots must possess semantic and visual perception abilities to transform real-world mappings of their workspace into the low-level control necessary for object manipulation. Recent work has demonstrated the capabilities of fine-tuning large Vision-Language Models (VLMs) to learn the mapping between RGB images, language instructions, and joint space control. These models typically take as input RGB images of the workspace and language instructions, and are trained on large datasets of teleoperated robot demonstrations. In this work, we explore methods to improve the scene context awareness of a popular recent Vision-Language-Action model by integrating chain-of-thought reasoning, depth perception, and task-oriented region of interest detection. Our experiments in the LIBERO simulation environment show that our proposed model, 3D-CAVLA, improves the success rate across various LIBERO task suites, achieving an average success rate of 98.1$\%$. We also evaluate the zero-shot capabilities of our method, demonstrating that 3D scene awareness leads to robust learning and adaptation for completely unseen tasks. 3D-CAVLA achieves an absolute improvement of 8.8$\%$ on unseen tasks. We will open-source our code and the unseen tasks dataset to promote community-driven research here: https://3d-cavla.github.io

📄 PDF Abstract BibTeX arXiv:2505.05800

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language-Action

Similar Papers 제목 키워드 기반

Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator

2025-02-26 · Xiankang He, Dongyan Guo, Hongji Li, Ruibo Li 외

Recent advances in zero-shot monocular depth estimation(MDE) have significantly improved generalization by unifying depth distributions through normalized depth representations and by leveraging large-scale unlabeled dat…

Depth EstimationDiversityMonocular Depth EstimationPseudo Label+1

GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation

2025-10-21 · Tuan Pham, Thanh-Tung Le, Xiaohui Xie, Stephan Mandt arxiv

We introduce a novel framework for metric depth estimation that enhances pretrained diffusion-based monocular depth estimation (DB-MDE) models with stereo vision guidance. While existing DB-MDE methods excel at predictin…

Monocular Depth Estimation

Fusion of stereo and still monocular depth estimates in a self-supervised learning context

2018-03-20 · Diogo Martins, Kevin van Hecke, Guido de Croon

We study how autonomous robots can learn by themselves to improve their depth estimation capability. In particular, we investigate a self-supervised learning setup in which stereo vision depth estimates serve as targets …

Autonomous NavigationDepth EstimationSelf-Supervised Learning

Semi-MoreGAN: A New Semi-supervised Generative Adversarial Network for Mixture of Rain Removal

2022-04-28 · Yiyang Shen, Yongzhen Wang, Mingqiang Wei, Honghua Chen 외

Rain is one of the most common weather which can completely degrade the image quality and interfere with the performance of many computer vision tasks, especially under heavy rain conditions. We observe that: (i) rain is…

Depth EstimationDepth PredictionGenerative Adversarial NetworkRain Removal

GEOcc: Geometrically Enhanced 3D Occupancy Network with Implicit-Explicit Depth Fusion and Contextual Self-Supervision

2024-05-17 · Xin Tan, Wenbin Wu, Zhiwei Zhang, Chaojie Fan 외

3D occupancy perception holds a pivotal role in recent vision-centric autonomous driving systems by converting surround-view images into integrated geometric and semantic representations within dense 3D grids. Neverthele…

Autonomous DrivingDecoderDepth EstimationDepth Prediction+1