paper-with-me

홈 › Papers

HRDFuse: Monocular 360°Depth Estimation by Collaboratively Learning Holistic-with-Regional Depth Distributions

2023-03-21 · Hao Ai, Zidong Cao, Yan-Pei Cao, Ying Shan, Lin Wang

Depth estimation from a monocular 360{\deg} image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, \eg, OmniFusion, have applied the tangent projection (TP) to represent a 360{\deg}image and predicted depth values via patch-wise regressions, which are merged to get a depth map with equirectangular projection (ERP) format. However, these methods suffer from 1) non-trivial process of merging plenty of patches; 2) capturing less holistic-with-regional contextual information by directly regressing the depth value of each pixel. In this paper, we propose a novel framework, \textbf{HRDFuse}, that subtly combines the potential of convolutional neural networks (CNNs) and transformers by collaboratively learning the \textit{holistic} contextual information from the ERP and the \textit{regional} structural information from the TP. Firstly, we propose a spatial feature alignment (\textbf{SFA}) module that learns feature similarities between the TP and ERP to aggregate the TP features into a complete ERP feature map in a pixel-wise manner. Secondly, we propose a collaborative depth distribution classification (\textbf{CDDC}) module that learns the \textbf{holistic-with-regional} histograms capturing the ERP and TP depth distributions. As such, the final depth values can be predicted as a linear combination of histogram bin centers. Lastly, we adaptively combine the depth predictions from ERP and TP to obtain the final depth map. Extensive experiments show that our method predicts\textbf{ more smooth and accurate depth} results while achieving \textbf{favorably better} results than the SOTA methods.

📄 PDF Abstract BibTeX arXiv:2303.11616

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationERP

Similar Papers 제목 키워드 기반

HRDFuse: Monocular 360deg Depth Estimation by Collaboratively Learning Holistic-With-Regional Depth Distributions

2023-01-01 · CVPR 2023 1 · Hao Ai, Zidong Cao, Yan-Pei Cao, Ying Shan 외

Depth estimation from a monocular 360 image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, e.g., OmniFusion, have applied the tangent projection (TP) to represent a 360 imag…

Depth EstimationERP

Self-supervised Monocular Depth and Pose Estimation for Endoscopy with Generative Latent Priors

2024-11-26 · Ziang Xu, Bin Li, Yang Hu, Chenyu Zhang 외

Accurate 3D mapping in endoscopy enables quantitative, holistic lesion characterization within the gastrointestinal (GI) tract, requiring reliable depth and pose estimation. However, endoscopy systems are monocular, and …

Pose Estimation

Hierarchical Normalization for Robust Monocular Depth Estimation

2022-10-18 · Chi Zhang, Wei Yin, Zhibin Wang, Gang Yu 외

In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-the-art methods adopt image-level normaliz…

Depth EstimationMonocular Depth Estimation

DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos

2024-03-09 · Xiuzhe Wu, Xiaoyang Lyu, Qihao Huang, Yong liu 외

Although considerable advancements have been attained in self-supervised depth estimation from monocular videos, most existing methods often treat all objects in a video as static entities, which however violates the dyn…

Depth EstimationDisentanglementMotion DisentanglementMotion Estimation+3

Distilled Semantics for Comprehensive Scene Understanding from Videos

2020-03-31 · CVPR 2020 6 · Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi 외

Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explic…

Depth EstimationKnowledge DistillationMonocular Depth EstimationMotion Segmentation+2