paper-with-me

Papers

HiMODE: A Hybrid Monocular Omnidirectional Depth Estimation Model

2022-04-11 · Masum Shah Junayed, Arezoo Sadeghzadeh, Md Baharul Islam, Lai-Kuan Wong, Tarkan Aydin

Monocular omnidirectional depth estimation is receiving considerable research attention due to its broad applications for sensing 360{\deg} surroundings. Existing approaches in this field suffer from limitations in recovering small object details and data lost during the ground-truth depth map acquisition. In this paper, a novel monocular omnidirectional depth estimation model, namely HiMODE is proposed based on a hybrid CNN+Transformer (encoder-decoder) architecture whose modules are efficiently designed to mitigate distortion and computational cost, without performance degradation. Firstly, we design a feature pyramid network based on the HNet block to extract high-resolution features near the edges. The performance is further improved, benefiting from a self and cross attention layer and spatial/temporal patches in the Transformer encoder and decoder, respectively. Besides, a spatial residual block is employed to reduce the number of parameters. By jointly passing the deep features extracted from an input image at each backbone block, along with the raw depth maps predicted by the transformer encoder-decoder, through a context adjustment layer, our model can produce resulting depth maps with better visual quality than the ground-truth. Comprehensive ablation studies demonstrate the significance of each individual module. Extensive experiments conducted on three datasets; Stanford3D, Matterport3D, and SunCG, demonstrate that HiMODE can achieve state-of-the-art performance for 360{\deg} monocular depth estimation.

📄 PDF Abstract BibTeX arXiv:2204.05007

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDepth EstimationmodelMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Batch Normalization 설명 없음

Similar Papers 제목 키워드 기반

PanoDepth: A Two-Stage Approach for Monocular Omnidirectional Depth Estimation

2022-02-02 · Yuyan Li, Zhixin Yan, Ye Duan, Liu Ren

Omnidirectional 3D information is essential for a wide range of applications such as Virtual Reality, Autonomous Driving, Robotics, etc. In this paper, we propose a novel, model-agnostic, two-stage pipeline for omnidirec…

Autonomous DrivingDepth EstimationMonocular Depth EstimationStereo Matching+1

Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model

2025-03-30 · Jannik Endres, Oliver Hahn, Charles Corbière, Simone Schaub-Meyer 외

Omnidirectional depth perception is essential for mobile robotics applications that require scene understanding across a full 360{\deg} field of view. Camera-based setups offer a cost-effective option by using stereo dep…

Depth EstimationMonocular Depth EstimationOmnnidirectional Stereo Depth EstimationScene Understanding+2

Distortion-Tolerant Monocular Depth Estimation On Omnidirectional Images Using Dual-cubemap

2022-03-18 · Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao 외

Estimating the depth of omnidirectional images is more challenging than that of normal field-of-view (NFoV) images because the varying distortion can significantly twist an object's shape. The existing methods suffer fro…

Depth EstimationMonocular Depth Estimation

FreDSNet: Joint Monocular Depth and Semantic Segmentation with Fast Fourier Convolutions

2022-10-04 · Bruno Berenguel-Baeta, Jesus Bermudez-Cameo, Jose J. Guerrero

In this work we present FreDSNet, a deep learning solution which obtains semantic 3D understanding of indoor environments from single panoramas. Omnidirectional images reveal task-specific advantages when addressing scen…

Depth EstimationMonocular Depth EstimationScene UnderstandingSegmentation+1

Distortion-aware Monocular Depth Estimation for Omnidirectional Images

2020-10-18 · Hong-Xiang Chen, Kunhong Li, Zhiheng Fu, Mengyi Liu 외

A main challenge for tasks on panorama lies in the distortion of objects among images. In this work, we propose a Distortion-Aware Monocular Omnidirectional (DAMO) dense depth estimation network to address this challenge…

Depth EstimationMonocular Depth Estimation