paper-with-me

홈 › Papers

Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking

2024-06-06 · Jiyao Zhang, Weiyao Huang, Bo Peng, Mingdong Wu, Fei Hu, Zijian Chen, Bo Zhao, Hao Dong

6D Object Pose Estimation is a crucial yet challenging task in computer vision, suffering from a significant lack of large-scale datasets. This scarcity impedes comprehensive evaluation of model performance, limiting research advancements. Furthermore, the restricted number of available instances or categories curtails its applications. To address these issues, this paper introduces Omni6DPose, a substantial dataset characterized by its diversity in object categories, large scale, and variety in object materials. Omni6DPose is divided into three main components: ROPE (Real 6D Object Pose Estimation Dataset), which includes 332K images annotated with over 1.5M annotations across 581 instances in 149 categories; SOPE(Simulated 6D Object Pose Estimation Dataset), consisting of 475K images created in a mixed reality setting with depth simulation, annotated with over 5M annotations across 4162 instances in the same 149 categories; and the manually aligned real scanned objects used in both ROPE and SOPE. Omni6DPose is inherently challenging due to the substantial variations and ambiguities. To address this challenge, we introduce GenPose++, an enhanced version of the SOTA category-level pose estimation framework, incorporating two pivotal improvements: Semantic-aware feature extraction and Clustering-based aggregation. Moreover, we provide a comprehensive benchmarking analysis to evaluate the performance of previous methods on this large-scale dataset in the realms of 6D object pose estimation and pose tracking.

📄 PDF Abstract BibTeX arXiv:2406.04316

Code (0)

등록된 구현이 없습니다.

Tasks

6D Pose Estimation using RGBBenchmarkingMixed RealityObjectPose EstimationPose Tracking

Similar Papers 제목 키워드 기반

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

HybridPose: 6D Object Pose Estimation under Hybrid Representations

2020-01-07 · CVPR 2020 6 · Chen Song, Jiaru Song, Qi-Xing Huang

We introduce HybridPose, a novel 6D object pose estimation approach. HybridPose utilizes a hybrid intermediate representation to express different geometric information in the input image, including keypoints, edge vecto…

6D Pose Estimation using RGBObjectPose Estimationregression

Towards Universal Vision-language Omni-supervised Segmentation

2023-03-12 · Bowen Dong, Jiaxi Gu, Jianhua Han, Hang Xu 외

Existing open-world universal segmentation approaches usually leverage CLIP and pre-computed proposal masks to treat open-world segmentation tasks as proposal classification. However, 1) these works cannot handle univers…

Instance Segmentationobject-detectionObject DetectionPanoptic Segmentation+3

RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base

2025-06-23 · Kuanning Wang, Yuqian Fu, Tianyu Wang, Yanwei Fu 외

Accurate 6D pose estimation is key for robotic manipulation, enabling precise object localization for tasks like grasping. We present RAG-6DPose, a retrieval-augmented approach that leverages 3D CAD models as a knowledge…

6D Pose EstimationObject LocalizationPose EstimationRAG+1

Occlusion Robust 3D Human Pose Estimation with StridedPoseGraphFormer and Data Augmentation

2023-04-24 · Soubarna Banik, Patricia Gschoßmann, Alejandro Mendoza Garcia, Alois Knoll

Occlusion is an omnipresent challenge in 3D human pose estimation (HPE). In spite of the large amount of research dedicated to 3D HPE, only a limited number of studies address the problem of occlusion explicitly. To fill…

3D Human Pose EstimationData AugmentationOcclusion HandlingPose Estimation