Auto-Rectify Network for Unsupervised Indoor Depth Estimation
Single-View depth estimation using the CNNs trained from unlabelled videos has shown significant promise. However, excellent results have mostly been obtained in street-scene driving scenarios, and such methods often fail in other settings, particularly indoor videos taken by handheld devices. In this work, we establish that the complex ego-motions exhibited in handheld settings are a critical obstacle for learning depth. Our fundamental analysis suggests that the rotation behaves as noise during training, as opposed to the translation (baseline) which provides supervision signals. To address the challenge, we propose a data pre-processing method that rectifies training images by removing their relative rotations for effective learning. The significantly improved performance validates our motivation. Towards end-to-end learning without requiring pre-processing, we propose an Auto-Rectify Network with novel loss functions, which can automatically learn to rectify images during training. Consequently, our results outperform the previous unsupervised SOTA method by a large margin on the challenging NYUv2 dataset. We also demonstrate the generalization of our trained model in ScanNet and Make3D, and the universality of our proposed learning method on 7-Scenes and KITTI datasets.
Code (1)
Tasks
Depth EstimationMonocular Depth EstimationSelf-Supervised LearningTranslationSimilar Papers 제목 키워드 기반
Indoor GeoNet: Weakly Supervised Hybrid Learning for Depth and Pose Estimation
Humans naturally perceive a 3D scene in front of them through accumulation of information obtained from multiple interconnected projections of the scene and by interpreting their correspondence. This phenomenon has inspi…
Camera Pose EstimationPose EstimationPLNet: Plane and Line Priors for Unsupervised Indoor Depth Estimation
Unsupervised learning of depth from indoor monocular videos is challenging as the artificial environment contains many textureless regions. Fortunately, the indoor scenes are full of specific structures, such as planes a…
Depth EstimationUnsupervised Monocular Depth Prediction for Indoor Continuous Video Streams
This paper studies unsupervised monocular depth prediction problem. Most of existing unsupervised depth prediction algorithms are developed for outdoor scenarios, while the depth prediction work in the indoor environment…
Depth EstimationDepth PredictionEnsemble LearningPose Estimation+1P$^{2}$Net: Patch-match and Plane-regularization for Unsupervised Indoor Depth Estimation
This paper tackles the unsupervised depth estimation task in indoor environments. The task is extremely challenging because of the vast areas of non-texture regions in these scenes. These areas could overwhelm the optimi…
Depth EstimationMonocular Depth EstimationSuperpixelsP²Net: Patch-match and Plane-regularization for Unsupervised Indoor Depth Estimation
This paper tackles the unsupervised depth estimation task in indoor environments. The task is extremely challenging because of the vast areas of non-texture regions in these scenes. These areas could overwhelm the optimi…
Depth EstimationSuperpixels