Multi-Frame Self-Supervised Depth Estimation with Multi-Scale Feature Fusion in Dynamic Scenes
Multi-frame methods improve monocular depth estimation over single-frame approaches by aggregating spatial-temporal information via feature matching. However, the spatial-temporal feature leads to accuracy degradation in dynamic scenes. To enhance the performance, recent methods tend to propose complex architectures for feature matching and dynamic scenes. In this paper, we show that a simple learning framework, together with designed feature augmentation, leads to superior performance. (1) A novel dynamic objects detecting method with geometry explainability is proposed. The detected dynamic objects are excluded during training, which guarantees the static environment assumption and relieves the accuracy degradation problem of the multi-frame depth estimation. (2) Multi-scale feature fusion is proposed for feature matching in the multi-frame depth network, which improves feature matching, especially between frames with large camera motion. (3) The robust knowledge distillation with a robust teacher network and reliability guarantee is proposed, which improves the multi-frame depth estimation without computation complexity increase during the test. The experiments show that our proposed methods achieve great performance improvement on the multi-frame depth estimation.
Code (0)
등록된 구현이 없습니다.
Tasks
Depth EstimationKnowledge DistillationMonocular Depth EstimationUnsupervised Monocular Depth EstimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
FusionDepth: Complement Self-Supervised Monocular Depth Estimation with Cost Volume
Multi-view stereo depth estimation based on cost volume usually works better than self-supervised monocular depth estimation except for moving objects and low-textured surfaces. So in this paper, we propose a multi-frame…
Depth EstimationMonocular Depth EstimationStereo Depth EstimationMAL: Motion-Aware Loss with Temporal and Distillation Hints for Self-Supervised Depth Estimation
Depth perception is crucial for a wide range of robotic applications. Multi-frame self-supervised depth estimation methods have gained research interest due to their ability to leverage large-scale, unlabeled real-world …
Depth EstimationMonocular Depth EstimationExploring the Mutual Influence between Self-Supervised Single-Frame and Multi-Frame Depth Estimation
Although both self-supervised single-frame and multi-frame depth estimation methods only require unlabeled monocular videos for training, the information they leverage varies because single-frame methods mainly rely on a…
Depth EstimationSelf-Supervised Ego-Motion Estimation Based on Multi-Layer Fusion of RGB and Inferred Depth
In existing self-supervised depth and ego-motion estimation methods, ego-motion estimation is usually limited to only leveraging RGB information. Recently, several methods have been proposed to further improve the accura…
Motion EstimationSelf-Supervised LearningEGA-Depth: Efficient Guided Attention for Self-Supervised Multi-Camera Depth Estimation
The ubiquitous multi-camera setup on modern autonomous vehicles provides an opportunity to construct surround-view depth. Existing methods, however, either perform independent monocular depth estimations on each camera o…
Autonomous DrivingAutonomous VehiclesDepth Estimation