paper-with-me

홈 › Papers

Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation

2026-03-29 · Rachit Agarwal, Abhishek Joshi, Sathish Chalasani, Woo Jin Kim arxiv

Object pose estimation is a fundamental task in 3D vision with applications in robotics, AR/VR, and scene understanding. We address the challenge of category-level 9-DoF pose estimation (6D pose + 3Dsize) from RGB-D input, without relying on CAD models during inference. Existing depth-only methods achieve strong results but ignore semantic cues from RGB, while many RGB-D fusion models underperform due to suboptimal cross-modal fusion that fails to align semantic RGB cues with 3D geometric representations. We propose DeMo-Pose, a hybrid architecture that fuses monocular semantic features with depth-based graph convolutional representations via a novel multimodal fusion strategy. To further improve geometric reasoning, we introduce a novel Mesh-Point Loss (MPL) that leverages mesh structure during training without adding inference overhead. Our approach achieves real-time inference and significantly improves over state-of-the-art methods across object categories, outperforming the strong GPV-Pose baseline by 3.2\% on 3D IoU and 11.1\% on pose accuracy on the REAL275 benchmark. The results highlight the effectiveness of depth-RGB fusion and geometry-aware learning, enabling robust category-level 3D pose estimation for real-world applications.

📄 PDF Abstract BibTeX arXiv:2603.27533

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding3D Pose Estimation

Similar Papers 제목 키워드 기반

FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators

2023-10-05 · Haiping Wang, YuAn Liu, Bing Wang, Yujing Sun 외

Matching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn…

Image to Point Cloud RegistrationMetric LearningPoint Cloud Registration

SRFNet: Monocular Depth Estimation with Fine-grained Structure via Spatial Reliability-oriented Fusion of Frames and Events

2023-09-22 · Tianbo Pan, Zidong Cao, Lin Wang

Monocular depth estimation is a crucial task to measure distance relative to a camera, which is important for applications, such as robot navigation and self-driving. Traditional frame-based methods suffer from performan…

Depth EstimationMonocular Depth EstimationRobot Navigation

Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth Estimation

2021-08-19 · ICCV 2021 10 · Hyunyoung Jung, Eunhyeok Park, Sungjoo Yoo

Self-supervised monocular depth estimation has been widely studied, owing to its practical importance and recent promising improvements. However, most works suffer from limited supervision of photometric consistency, esp…

Depth EstimationMetric LearningMonocular Depth Estimation

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

2025-07-26 · Luoxi Jing, Dianxi Shi, Zhe Liu, Songchang Jin 외 arxiv

Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range…

Monocular Depth EstimationScene Understanding

Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

2024-02-19 · Jialei Xu, Xianming Liu, Junjun Jiang, Kui Jiang 외

Monocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared ca…

Depth EstimationMonocular Depth Estimation