paper-with-me

홈 › Papers

Mask6D: Masked Pose Priors For 6D Object Pose Estimation

2025-07-09 · Yuechen Xie, Haobo Jiang, Jin Xie arxiv

Robust 6D object pose estimation in cluttered or occluded conditions using monocular RGB images remains a challenging task. One reason is that current pose estimation networks struggle to extract discriminative, pose-aware features using 2D feature backbones, especially when the available RGB information is limited due to target occlusion in cluttered scenes. To mitigate this, we propose a novel pose estimation-specific pre-training strategy named Mask6D. Our approach incorporates pose-aware 2D-3D correspondence maps and visible mask maps as additional modal information, which is combined with RGB images for the reconstruction-based model pre-training. Essentially, this 2D-3D correspondence maps a transformed 3D object model to 2D pixels, reflecting the pose information of the target in camera coordinate system. Meanwhile, the integrated visible mask map can effectively guide our model to disregard cluttered background information. In addition, an object-focused pre-training loss function is designed to further facilitate our network to remove the background interference. Finally, we fine-tune our pre-trained pose prior-aware network via conventional pose training strategy to realize the reliable pose prediction. Extensive experiments verify that our method outperforms previous end-to-end pose estimation methods.

📄 PDF Abstract BibTeX arXiv:2507.06486

Code (0)

등록된 구현이 없습니다.

Tasks

Pose EstimationPose Prediction

Similar Papers 제목 키워드 기반

SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification

2025-06-26 · Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

Retrieving 3D objects in complex indoor environments using only a masked 2D image and a natural language description presents significant challenges. The ROOMELSA challenge limits access to full 3D scene context, complic…

3D Object RetrievalObjectRe-RankingRetrieval

Masked Visual Actions for Unified World Modeling

2026-07-21 · Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey 외 hf

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such…

Decision Making

PRISM: Progressive Restoration for Scene Graph-based Image Manipulation

2023-11-03 · Pavel Jahoda, Azade Farshad, Yousef Yeganeh, Ehsan Adeli 외

Scene graphs have emerged as accurate descriptive priors for image generation and manipulation tasks, however, their complexity and diversity of the shapes and relations of objects in data make it challenging to incorpor…

DenoisingDescriptiveDiversityImage Generation+1

MARMOT: Masked Autoencoder for Modeling Transient Imaging

2025-06-10 · Siyuan Shen, Ziheng Wang, Xingyue Peng, Suan Xia 외

Pretrained models have demonstrated impressive success in many modalities such as language and vision. Recent works facilitate the pretraining paradigm in imaging research. Transients are a novel modality, which are capt…

Decoder

M$^{3}$3D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding

2023-09-26 · Muhammad Abdullah Jamal, Omid Mohareri

We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learn…

2D Semantic SegmentationAction DetectionAction RecognitionContrastive Learning+7