paper-with-me

홈 › Papers

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

2025-06-14 · Wenbo Li, Shiyi Wang, Yiteng Chen, Huiping Zhuang, Qingyao Wu

Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediate representations, discarding important task-specific information such as fine-grained spatial or semantic details. To address this, we propose AntiGrounding, a new framework that reverses the instruction grounding process. It lifts candidate actions directly into the VLM representation space, renders trajectories from multiple views, and uses structured visual question answering for instruction-based decision making. This enables zero-shot synthesis of optimal closed-loop robot trajectories for new tasks. We also propose an offline policy refinement module that leverages past experience to enhance long-term performance. Experiments in both simulation and real-world environments show that our method outperforms baselines across diverse robotic manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2506.12374

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

Pre-training Auto-regressive Robotic Models with 4D Representations

2025-02-18 · Dantong Niu, Yuvan Sharma, Haoru Xue, Giscard Biamby 외

Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Y…

Depth EstimationMonocular Depth EstimationPoint TrackingTransfer Learning

Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation

2025-01-01 · CVPR 2025 1 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu 외

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has…

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

2024-11-27 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu 외

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has inc…

Multi-Object Grasping -- Estimating the Number of Objects in a Robotic Grasp

2021-11-30 · Tianze Chen, Adheesh Shenoy, Anzhelika Kolinko, Syed Shah 외

A human hand can grasp a desired number of objects at once from a pile based solely on tactile sensing. To do so, a robot needs to grasp within a pile, sense the number of objects in the grasp before lifting, and predict…

3D Human Pose Estimation in RGBD Images for Robotic Task Learning

2018-03-07 · Christian Zimmermann, Tim Welschehold, Christian Dornhege, Wolfram Burgard 외

We propose an approach to estimate 3D human pose in real world units from a single RGBD image and show that it exceeds performance of monocular 3D pose estimation approaches from color as well as pose estimation exclusiv…

3D Human Pose Estimation3D Pose EstimationPose Estimation