SA6D: Self-Adaptive Few-Shot 6D Pose Estimator for Novel and Occluded Objects
To enable meaningful robotic manipulation of objects in the real-world, 6D pose estimation is one of the critical aspects. Most existing approaches have difficulties to extend predictions to scenarios where novel object instances are continuously introduced, especially with heavy occlusions. In this work, we propose a few-shot pose estimation (FSPE) approach called SA6D, which uses a self-adaptive segmentation module to identify the novel target object and construct a point cloud model of the target object using only a small number of cluttered reference images. Unlike existing methods, SA6D does not require object-centric reference images or any additional object information, making it a more generalizable and scalable solution across categories. We evaluate SA6D on real-world tabletop object datasets and demonstrate that SA6D outperforms existing FSPE methods, particularly in cluttered scenes with occlusions, while requiring fewer reference images.
Code (0)
등록된 구현이 없습니다.
Tasks
6D Pose EstimationObjectPose EstimationSimilar Papers 제목 키워드 기반
Feature Calibration Network for Occluded Pedestrian Detection
Pedestrian detection in the wild remains a challenging problem especially for scenes containing serious occlusion. In this paper, we propose a novel feature learning method in the deep learning framework, referred to as …
Pedestrian DetectionExploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
To integrate action recognition into autonomous robotic systems, it is essential to address challenges such as person occlusions-a common yet often overlooked scenario in existing self-supervised skeleton-based action re…
Action RecognitionImputationSelf-Supervised LearningSelf-supervised Skeleton-based Action Recognition+2MHSA-Net: Multi-Head Self-Attention Network for Occluded Person Re-Identification
This paper presents a novel person re-identification model, named Multi-Head Self-Attention Network (MHSA-Net), to prune unimportant information and capture key local information from person images. MHSA-Net contains two…
DiversityOccluded Person Re-IdentificationPerson Re-IdentificationFast Self-Supervised depth and mask aware Association for Multi-Object Tracking
Multi-object tracking (MOT) methods often rely on Intersection-over-Union (IoU) for association. However, this becomes unreliable when objects are similar or occluded. Also, computing IoU for segmentation masks is comput…
Multi-Object TrackingDomes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction
Existing state-of-the-art estimation systems can detect 2d poses of multiple people in images quite reliably. In contrast, 3d pose estimation from a single image is ill-posed due to occlusion and depth ambiguities. Assum…
2D Pose Estimation3D Pose Estimation3D ReconstructionDeep Reinforcement Learning+2