Multiple Prior Representation Learning for Self-Supervised Monocular Depth Estimation via Hybrid Transformer
Self-supervised monocular depth estimation aims to infer depth information without relying on labeled data. However, the lack of labeled information poses a significant challenge to the model's representation, limiting its ability to capture the intricate details of the scene accurately. Prior information can potentially mitigate this issue, enhancing the model's understanding of scene structure and texture. Nevertheless, solely relying on a single type of prior information often falls short when dealing with complex scenes, necessitating improvements in generalization performance. To address these challenges, we introduce a novel self-supervised monocular depth estimation model that leverages multiple priors to bolster representation capabilities across spatial, context, and semantic dimensions. Specifically, we employ a hybrid transformer and a lightweight pose network to obtain long-range spatial priors in the spatial dimension. Then, the context prior attention is designed to improve generalization, particularly in complex structures or untextured areas. In addition, semantic priors are introduced by leveraging semantic boundary loss, and semantic prior attention is supplemented, further refining the semantic features extracted by the decoder. Experiments on three diverse datasets demonstrate the effectiveness of the proposed model. It integrates multiple priors to comprehensively enhance the representation ability, improving the accuracy and reliability of depth estimation. Codes are available at: \url{https://github.com/MVME-HBUT/MPRLNet}
Code (1)
Tasks
DecoderDepth EstimationMonocular Depth EstimationRepresentation LearningSimilar Papers 제목 키워드 기반
JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth…
Monocular Depth EstimationRepresentation LearningMonocular Depth Estimation with Self-supervised Instance Adaptation
Recent advances in self-supervised learning havedemonstrated that it is possible to learn accurate monoculardepth reconstruction from raw video data, without using any 3Dground truth for supervision. However, in robotics…
Depth EstimationMonocular Depth EstimationMonocular ReconstructionSelf-Supervised LearningDINOcular: Self-Supervised Visuospatial Representations
We introduce a self-supervised framework for learning joint visuospatial representations from RGB-D observations. While modern vision foundation models are trained almost exclusively on RGB images, many embodied systems …
Semantic SegmentationEnhancing self-supervised monocular depth estimation with traditional visual odometry
Estimating depth from a single image represents an attractive alternative to more traditional approaches leveraging multiple cameras. In this field, deep learning yielded outstanding results at the cost of needing large …
Depth And Camera MotionDepth EstimationMonocular Depth EstimationVisual OdometryTime-to-Label: Temporal Consistency for Self-Supervised Monocular 3D Object Detection
Monocular 3D object detection continues to attract attention due to the cost benefits and wider availability of RGB cameras. Despite the recent advances and the ability to acquire data at scale, annotation cost and compl…
3D Object DetectionDepth EstimationMonocular 3D Object DetectionObject+2