IEBins: Iterative Elastic Bins for Monocular Depth Estimation
Monocular depth estimation (MDE) is a fundamental topic of geometric computer vision and a core technique for many downstream applications. Recently, several methods reframe the MDE as a classification-regression problem where a linear combination of probabilistic distribution and bin centers is used to predict depth. In this paper, we propose a novel concept of iterative elastic bins (IEBins) for the classification-regression-based MDE. The proposed IEBins aims to search for high-quality depth by progressively optimizing the search range, which involves multiple stages and each stage performs a finer-grained depth search in the target bin on top of its previous stage. To alleviate the possible error accumulation during the iterative process, we utilize a novel elastic target bin to replace the original target bin, the width of which is adjusted elastically based on the depth uncertainty. Furthermore, we develop a dedicated framework composed of a feature extractor and an iterative optimizer that has powerful temporal context modeling capabilities benefiting from the GRU-based architecture. Extensive experiments on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets demonstrate that the proposed method surpasses prior state-of-the-art competitors. The source code is publicly available at https://github.com/ShuweiShao/IEBins.
Code (1)
Tasks
Depth EstimationMonocular Depth EstimationregressionSimilar Papers 제목 키워드 기반
BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation
Monocular depth estimation is a fundamental task in computer vision and has drawn increasing attention. Recently, some methods reformulate it as a classification-regression task to boost the model performance, where cont…
DecoderDepth EstimationMonocular Depth Estimationregression+1Depthformer : Multiscale Vision Transformer For Monocular Depth Estimation With Local Global Information Fusion
Attention-based models such as transformers have shown outstanding performance on dense prediction tasks, such as semantic segmentation, owing to their capability of capturing long-range dependency in an image. However, …
DecoderDepth EstimationDepth PredictionMonocular Depth Estimation+1Adaptive Discrete Disparity Volume for Self-supervised Monocular Depth Estimation
In self-supervised monocular depth estimation tasks, discrete disparity prediction has been proven to attain higher quality depth maps than common continuous methods. However, current discretization strategies often divi…
Depth EstimationMonocular Depth EstimationEstimating Depth from Monocular Images as Classification Using Deep Fully Convolutional Residual Networks
Depth estimation from single monocular images is a key component of scene understanding and has benefited largely from deep convolutional neural networks (CNN) recently. In this article, we take advantage of the recent d…
Depth EstimationGeneral ClassificationScene UnderstandingLearning to Adapt CLIP for Few-Shot Monocular Depth Estimation
Pre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities. When CLIP is used for depth estimation ta…
Depth EstimationMonocular Depth Estimation