paper-with-me

홈 › Papers

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

2024-11-27 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu, Zhilue Wang, Longzan Luo, Lily Lee, Pengwei Wang, Zhongyuan Wang, Renrui Zhang, Shanghang Zhang

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has increasingly focused on the explicit extraction of 3D features, while still facing challenges such as the lack of large-scale robotic 3D data and the potential loss of spatial geometry. To address these limitations, we propose the Lift3D framework, which progressively enhances 2D foundation models with implicit and explicit 3D robotic representations to construct a robust 3D manipulation policy. Specifically, we first design a task-aware masked autoencoder that masks task-relevant affordance patches and reconstructs depth information, enhancing the 2D foundation model's implicit 3D robotic representation. After self-supervised fine-tuning, we introduce a 2D model-lifting strategy that establishes a positional mapping between the input 3D points and the positional embeddings of the 2D model. Based on the mapping, Lift3D utilizes the 2D foundation model to directly encode point cloud data, leveraging large-scale pretrained knowledge to construct explicit 3D robotic representations while minimizing spatial information loss. In experiments, Lift3D consistently outperforms previous state-of-the-art methods across several simulation benchmarks and real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2411.18623

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation

2025-01-01 · CVPR 2025 1 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu 외

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has…

AugLift: Depth-Aware Input Reparameterization Improves Domain Generalization in 2D-to-3D Pose Lifting

2025-08-09 · Nikolai Warner, Wenjin Zhang, Hamid Badiozamani, Irfan Essa 외 arxiv

Lifting-based 3D human pose estimation infers 3D joints from 2D keypoints but generalizes poorly because $(x,y)$ coordinates alone are an ill-posed, sparse representation that discards geometric information modern founda…

3D Human Pose EstimationDomain Generalization

PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters

2026-03-04 · Yinghong Yu, Guangyuan Li, Jiancheng Yang arxiv

Large-scale 2D foundation models exhibit strong transferable representations, yet extending them to 3D volumetric data typically requires retraining, adapters, or architectural redesign. We introduce PlaneCycle, a traini…

3D Classification

NeuroLifting: Neural Inference on Markov Random Fields at Scale

2024-11-28 · Yaomin Wang, Chaolong Ying, Xiaodong Luo, Tianshu Yu

Inference in large-scale Markov Random Fields (MRFs) is a critical yet challenging task, traditionally approached through approximate methods like belief propagation and mean field, or exact methods such as the Toulbar2 …

Residual Reinforcement Learning for Waste-Container Lifting Using Large-Scale Cranes with Underactuated Tools

2026-02-05 · Qi Li, Karsten Berns arxiv

This paper studies the container lifting phase of a waste-container recycling task in urban environments, performed by a hydraulic loader crane equipped with an underactuated discharge unit, and proposes a residual reinf…

Reinforcement Learning