Two Stream Networks for Self-Supervised Ego-Motion Estimation
Learning depth and camera ego-motion from raw unlabeled RGB video streams is seeing exciting progress through self-supervision from strong geometric cues. To leverage not only appearance but also scene geometry, we propose a novel self-supervised two-stream network using RGB and inferred depth information for accurate visual odometry. In addition, we introduce a sparsity-inducing data augmentation policy for ego-motion learning that effectively regularizes the pose network to enable stronger generalization performance. As a result, we show that our proposed two-stream pose network achieves state-of-the-art results among learning-based methods on the KITTI odometry benchmark, and is especially suited for self-supervision at scale. Our experiments on a large-scale urban driving dataset of 1 million frames indicate that the performance of our proposed architecture does indeed scale progressively with more data.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationMotion EstimationVisual OdometryVocal Bursts Valence PredictionSimilar Papers 제목 키워드 기반
Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals
Estimating motion in videos is an essential computer vision problem with many downstream applications, including controllable video generation and robotics. Current solutions are primarily trained using synthetic data or…
counterfactualMotion EstimationOcclusion EstimationSelf-Supervised Learning+1MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features
Self-supervised learning of visual representations has been focusing on learning content features, which do not capture object motion or location, and focus on identifying and differentiating objects in images and videos…
Optical Flow EstimationSelf-Supervised LearningSemantic SegmentationDo not trust the neighbors! Adversarial Metric Learning for Self-Supervised Scene Flow Estimation
Scene flow is the task of estimating 3D motion vectors to individual points of a dynamic 3D scene. Motion vectors have shown to be beneficial for downstream tasks such as action classification and collision avoidance. Ho…
Action ClassificationCollision AvoidanceMetric LearningScene Flow Estimation+2Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?
Geometric constraints are shown to enforce scale consistency and remedy the scale ambiguity issue in self-supervised monocular depth estimation. Meanwhile, scale-invariant losses focus on learning relative depth, lea…
Depth EstimationDepth PredictionMonocular Depth EstimationMotion EstimationSelf-Supervised Ego-Motion Estimation Based on Multi-Layer Fusion of RGB and Inferred Depth
In existing self-supervised depth and ego-motion estimation methods, ego-motion estimation is usually limited to only leveraging RGB information. Recently, several methods have been proposed to further improve the accura…
Motion EstimationSelf-Supervised Learning