paper-with-me

홈 › Papers

POPE: 6-DoF Promptable Pose Estimation of Any Object, in Any Scene, with One Reference

2023-05-25 · Zhiwen Fan, Panwang Pan, Peihao Wang, Yifan Jiang, Dejia Xu, Hanwen Jiang, Zhangyang Wang

Despite the significant progress in six degrees-of-freedom (6DoF) object pose estimation, existing methods have limited applicability in real-world scenarios involving embodied agents and downstream 3D vision tasks. These limitations mainly come from the necessity of 3D models, closed-category detection, and a large number of densely annotated support views. To mitigate this issue, we propose a general paradigm for object pose estimation, called Promptable Object Pose Estimation (POPE). The proposed approach POPE enables zero-shot 6DoF object pose estimation for any target object in any scene, while only a single reference is adopted as the support view. To achieve this, POPE leverages the power of the pre-trained large-scale 2D foundation model, employs a framework with hierarchical feature representation and 3D geometry principles. Moreover, it estimates the relative camera pose between object prompts and the target object in new views, enabling both two-view and multi-view 6DoF pose estimation tasks. Comprehensive experimental results demonstrate that POPE exhibits unrivaled robust performance in zero-shot settings, by achieving a significant reduction in the averaged Median Pose Error by 52.38% and 50.47% on the LINEMOD and OnePose datasets, respectively. We also conduct more challenging testings in causally captured images (see Figure 1), which further demonstrates the robustness of POPE. Project page can be found with https://paulpanwang.github.io/POPE/.

📄 PDF Abstract BibTeX arXiv:2305.15727

Code (1)

paulpanwang/POPE pytorch

Tasks

3D geometryObjectPose Estimation

Similar Papers 제목 키워드 기반

GMS-VINS:Multi-category Dynamic Objects Semantic Segmentation for Enhanced Visual-Inertial Odometry Using a Promptable Foundation Model

2024-11-28 · Rui Zhou, Jingbin Liu, Junbin Xie, Jianyu Zhang 외

Visual-inertial odometry (VIO) is widely used in various fields, such as robots, drones, and autonomous vehicles, due to its low cost and complementary sensors. Most VIO methods presuppose that observed objects are stati…

Autonomous VehiclesPose EstimationSemantic Segmentation

PromptHMR: Promptable Human Mesh Recovery

2025-04-08 · CVPR 2025 1 · Yufu Wang, Yu Sun, Priyanka Patel, Kostas Daniilidis 외

Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxili…

3D Human Pose EstimationHuman Mesh Recovery

SRPose: Two-view Relative Pose Estimation with Sparse Keypoints

2024-07-11 · Rui Yin, Yulun Zhang, Zherong Pan, Jianjun Zhu 외

Two-view pose estimation is essential for map-free visual relocalization and object pose tracking tasks. However, traditional matching methods suffer from time-consuming robust estimators, while deep learning-based pose …

ObjectPose EstimationPose Tracking

SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

2024-08-29 · Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Chengzhuo Tong 외

We introduce SAM2Point, a preliminary exploration adapting Segment Anything Model 2 (SAM 2) for zero-shot and promptable 3D segmentation. SAM2Point interprets any 3D data as a series of multi-directional videos, and leve…

Segmentation

Scene and Human in One World: Reconstruction in a Feedforward Pass

2026-06-26 · Boao Shi, Qiao Feng, Yiming Huang, Lingjie Liu arxiv

Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusion interference. Rather than treating human mesh recovery and scene r…

Human Mesh Recovery