paper-with-me

홈 › Papers

Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation

2025-01-01 · CVPR 2025 1 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu, Zhilve Wang, Longzan Luo, Xiaoqi Li, Pengwei Wang, Zhongyuan Wang, Renrui Zhang, Shanghang Zhang

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has increasingly focused on the explicit extraction of 3D features, while still facing challenges such as the lack of large-scale robotic 3D data and the potential loss of spatial geometry. To address these limitations, we propose the Lift3D framework, which progressively enhances 2D foundation models with implicit and explicit 3D robotic representations to construct a robust 3D manipulation policy. Specifically, we first design a task-aware masked autoencoder that masks task-relevant affordance patches and reconstructs depth information, enhancing the 2D foundation model's implicit 3D robotic representation. After self-supervised fine-tuning, we introduce a 2D model-lifting strategy that establishes a positional mapping between the input 3D points and the positional embeddings of the 2D model. Based on the mapping, Lift3D utilizes the 2D foundation model to directly encode point cloud data, leveraging large-scale pretrained knowledge to construct explicit 3D robotic representations while minimizing spatial information loss. In experiments, Lift3D consistently outperforms previous state-of-the-art methods across several simulation benchmarks and real-world scenarios.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

2024-11-27 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu 외

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has inc…

AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation

2026-02-16 · Morgan Byrd, Donghoon Baek, Kartik Garg, Hyunyoung Jung 외 arxiv

This paper presents Adaptive Whole-body Loco-Manipulation, AdaptManip, a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery. Unlike prior imitation learning-base…

Reinforcement Learning

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

2025-06-14 · Wenbo Li, Shiyi Wang, Yiteng Chen, Huiping Zhuang 외

Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediat…

Decision MakingQuestion AnsweringVisual Question Answering

Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon Tasks

2025-07-06 · Hao Huang, Shuaihang Yuan, Geeta Chandra Raju Bethala, Congcong Wen 외 arxiv

Policy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves hand…

Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

2026-06-30 · Zixing Wang, Kausik Sivakumar, Jinghuan Shang, Yafei Hu 외 arxiv

Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet…

Motion Planning