paper-with-me

홈 › Papers

BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

2022-04-03 · Zhenyu Li, Xuyang Wang, Xianming Liu, Junjun Jiang

Monocular depth estimation is a fundamental task in computer vision and has drawn increasing attention. Recently, some methods reformulate it as a classification-regression task to boost the model performance, where continuous depth is estimated via a linear combination of predicted probability distributions and discrete bins. In this paper, we present a novel framework called BinsFormer, tailored for the classification-regression-based depth estimation. It mainly focuses on two crucial components in the specific task: 1) proper generation of adaptive bins and 2) sufficient interaction between probability distribution and bins predictions. To specify, we employ the Transformer decoder to generate bins, novelly viewing it as a direct set-to-set prediction problem. We further integrate a multi-scale decoder structure to achieve a comprehensive understanding of spatial geometry information and estimate depth maps in a coarse-to-fine manner. Moreover, an extra scene understanding query is proposed to improve the estimation accuracy, which turns out that models can implicitly learn useful information from an auxiliary environment classification task. Extensive experiments on the KITTI, NYU, and SUN RGB-D datasets demonstrate that BinsFormer surpasses state-of-the-art monocular depth estimation methods with prominent margins. Code and pretrained models will be made publicly available at \url{https://github.com/zhyever/Monocular-Depth-Estimation-Toolbox}.

📄 PDF Abstract BibTeX arXiv:2204.00987

Code (2)

zhyever/monocular-depth-estimation-toolbox 공식 구현 pytorch
RuijieZhu94/mmdepth pytorch

Tasks

DecoderDepth EstimationMonocular Depth EstimationregressionScene Understanding

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Adaptive Discrete Disparity Volume for Self-supervised Monocular Depth Estimation

2024-04-04 · Jianwei Ren

In self-supervised monocular depth estimation tasks, discrete disparity prediction has been proven to attain higher quality depth maps than common continuous methods. However, current discretization strategies often divi…

Depth EstimationMonocular Depth Estimation

Depthformer : Multiscale Vision Transformer For Monocular Depth Estimation With Local Global Information Fusion

2022-07-10 · Ashutosh Agarwal, Chetan Arora

Attention-based models such as transformers have shown outstanding performance on dense prediction tasks, such as semantic segmentation, owing to their capability of capturing long-range dependency in an image. However, …

DecoderDepth EstimationDepth PredictionMonocular Depth Estimation+1

A Vanilla Multi-Task Framework for Dense Visual Prediction Solution to 1st VCL Challenge -- Multi-Task Robustness Track

2024-02-27 · Zehui Chen, Qiuchen Wang, Zhenyu Li, Jiaming Liu 외

In this report, we present our solution to the multi-task robustness track of the 1st Visual Continual Learning (VCL) Challenge at ICCV 2023 Workshop. We propose a vanilla framework named UniNet that seamlessly combines …

3D Object DetectionContinual LearningDepth EstimationInstance Segmentation+4

IEBins: Iterative Elastic Bins for Monocular Depth Estimation

2023-09-25 · NeurIPS 2023 11 · Shuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu 외

Monocular depth estimation (MDE) is a fundamental topic of geometric computer vision and a core technique for many downstream applications. Recently, several methods reframe the MDE as a classification-regression problem…

Depth EstimationMonocular Depth Estimationregression

SlaBins: Fisheye Depth Estimation using Slanted Bins on Road Environments

2023-01-01 · ICCV 2023 1 · Jongsung Lee, Gyeongsu Cho, Jeongin Park, Kyongjun Kim 외

Although 3D perception for autonomous vehicles has focused on frontal-view information, more than half of fatal accidents occur due to side impacts in practice (e.g., T-bone crash). Motivated by this fact, we investi…

Autonomous VehiclesDepth Estimation