Unsupervised Scale-consistent Depth Learning from Video
We propose a monocular depth estimator SC-Depth, which requires only unlabelled videos for training and enables the scale-consistent prediction at inference time. Our contributions include: (i) we propose a geometry consistency loss, which penalizes the inconsistency of predicted depths between adjacent views; (ii) we propose a self-discovered mask to automatically localize moving objects that violate the underlying static scene assumption and cause noisy signals during training; (iii) we demonstrate the efficacy of each component with a detailed ablation study and show high-quality depth estimation results in both KITTI and NYUv2 datasets. Moreover, thanks to the capability of scale-consistent prediction, we show that our monocular-trained deep networks are readily integrated into the ORB-SLAM2 system for more robust and accurate tracking. The proposed hybrid Pseudo-RGBD SLAM shows compelling results in KITTI, and it generalizes well to the KAIST dataset without additional training. Finally, we provide several demos for qualitative evaluation.
Code (2)
Tasks
Depth EstimationMonocular Depth EstimationMonocular Visual OdometrySimultaneous Localization and MappingSimilar Papers 제목 키워드 기반
Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static…
Camera Pose EstimationDepth And Camera MotionDepth EstimationMonocular Depth Estimation+1DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth Completion
Unsupervised depth completion aims to recover dense depth from the sparse one without using the ground-truth annotation. Although depth measurement obtained from LiDAR is usually sparse, it contains valid and real distan…
Depth CompletionDepth EstimationDepth PredictionvalidDepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
Diffusion-based video depth estimation methods have achieved remarkable success with strong generalization ability. However, predicting depth for long videos remains challenging. Existing methods typically split videos i…
Depth EstimationScene-Centric Unsupervised Video Panoptic Segmentation
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any …
Video Panoptic SegmentationScene UnderstandingVideo SegmentationImage SegmentationUnsupervised Learning of Depth and Ego-Motion from Monocular Video Using 3D Geometric Constraints
We present a novel approach for unsupervised learning of depth and ego-motion from monocular video. Unsupervised learning removes the need for separate supervisory signals (depth or ego-motion ground truth, or multi-view…
3D geometryDepth And Camera MotionDepth EstimationMonocular Depth Estimation+1