Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?
Geometric constraints are shown to enforce scale consistency and remedy the scale ambiguity issue in self-supervised monocular depth estimation. Meanwhile, scale-invariant losses focus on learning relative depth, leading to accurate relative depth prediction. To combine the best of both worlds, we learn scale-consistent self-supervised depth in a scale-invariant manner. Towards this goal, we present a scale-aware geometric (SAG) loss, which enforces scale consistency through point cloud alignment. Compared to prior arts, SAG loss takes relative scale into consideration during relative motion estimation, enabling more precise alignment and explicit supervision for scale inference. In addition, a novel two-stream architecture for depth estimation is designed, which disentangles scale from depth estimation and allows depth to be learned in a scale-invariant manner. The integration of SAG loss and two-stream network enables more consistent scale inference and more accurate relative depth estimation. Our method achieves state-of-the-art performance under both scale-invariant and scale-dependent evaluation settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Depth EstimationDepth PredictionMonocular Depth EstimationMotion EstimationSimilar Papers 제목 키워드 기반
Multimodal Scale Consistency and Awareness for Monocular Self-Supervised Depth Estimation
Dense depth estimation is essential to scene-understanding for autonomous driving. However, recent self-supervised approaches on monocular videos suffer from scale-inconsistency across long sequences. Utilizing data from…
Autonomous DrivingDepth EstimationMonocular Depth EstimationScene UnderstandingPose Constraints for Consistent Self-supervised Monocular Depth and Ego-motion
Self-supervised monocular depth estimation approaches suffer not only from scale ambiguity but also infer temporally inconsistent depth maps w.r.t. scale. While disambiguating scale during training is not possible withou…
Camera Pose EstimationDepth EstimationEgocentric Pose EstimationMonocular Depth Estimation+2Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static…
Camera Pose EstimationDepth And Camera MotionDepth EstimationMonocular Depth Estimation+1Monocular Depth Estimation with Self-supervised Instance Adaptation
Recent advances in self-supervised learning havedemonstrated that it is possible to learn accurate monoculardepth reconstruction from raw video data, without using any 3Dground truth for supervision. However, in robotics…
Depth EstimationMonocular Depth EstimationMonocular ReconstructionSelf-Supervised LearningAVS-Net: Audio-Visual Scale Net for Self-supervised Monocular Metric Depth Estimation
Metric depth prediction from monocular videos suffers from bad generalization between datasets and requires supervised depth data for scale-correct training. Self-supervised training using multi-view reconstruction can b…
Depth EstimationDepth PredictionPrediction