paper-with-me

홈 › Papers

3D-LFM: Lifting Foundation Model

2023-12-19 · CVPR 2024 1 · Mosam Dabhi, Laszlo A. Jeni, Simon Lucey

The lifting of 3D structure and camera from 2D landmarks is at the cornerstone of the entire discipline of computer vision. Traditional methods have been confined to specific rigid objects, such as those in Perspective-n-Point (PnP) problems, but deep learning has expanded our capability to reconstruct a wide range of object classes (e.g. C3DPO and PAUL) with resilience to noise, occlusions, and perspective distortions. All these techniques, however, have been limited by the fundamental need to establish correspondences across the 3D training data -- significantly limiting their utility to applications where one has an abundance of "in-correspondence" 3D data. Our approach harnesses the inherent permutation equivariance of transformers to manage varying number of points per 3D data instance, withstands occlusions, and generalizes to unseen categories. We demonstrate state of the art performance across 2D-3D lifting task benchmarks. Since our approach can be trained across such a broad class of structures we refer to it simply as a 3D Lifting Foundation Model (3D-LFM) -- the first of its kind.

📄 PDF Abstract BibTeX arXiv:2312.11894

Code (1)

mosamdabhi/3dlfm 공식 구현

Tasks

3D Facial Landmark Localization3D Hand Pose Estimation3D Human Pose EstimationmodelMonocular 3D Human Pose EstimationPose Estimation

Similar Papers 제목 키워드 기반

PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters

2026-03-04 · Yinghong Yu, Guangyuan Li, Jiancheng Yang arxiv

Large-scale 2D foundation models exhibit strong transferable representations, yet extending them to 3D volumetric data typically requires retraining, adapters, or architectural redesign. We introduce PlaneCycle, a traini…

3D Classification

AugLift: Depth-Aware Input Reparameterization Improves Domain Generalization in 2D-to-3D Pose Lifting

2025-08-09 · Nikolai Warner, Wenjin Zhang, Hamid Badiozamani, Irfan Essa 외 arxiv

Lifting-based 3D human pose estimation infers 3D joints from 2D keypoints but generalizes poorly because $(x,y)$ coordinates alone are an ill-posed, sparse representation that discards geometric information modern founda…

3D Human Pose EstimationDomain Generalization

Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation

2025-01-01 · CVPR 2025 1 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu 외

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has…

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

2024-11-27 · Yueru Jia, Jiaming Liu, Sixiang Chen, Chenyang Gu 외

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has inc…

EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

2024-06-14 · Julian Straub, Daniel DeTone, Tianwei Shen, Nan Yang 외

The advent of wearable computers enables a new source of context for AI that is embedded in egocentric sensor data. This new egocentric data comes equipped with fine-grained 3D location information and thus presents the …

3D Object Detection3D ReconstructionMulti-View 3D Reconstructionobject-detection+1