Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
We present GLNet, a self-supervised framework for learning depth, optical flow, camera pose and intrinsic parameters from monocular video - addressing the difficulty of acquiring realistic ground-truth for such tasks. We propose three contributions: 1) we design new loss functions that capture multiple geometric constraints (eg. epipolar geometry) as well as an adaptive photometric loss that supports multiple moving objects, rigid and non-rigid, 2) we extend the model such that it predicts camera intrinsics, making it applicable to uncalibrated video, and 3) we propose several online refinement strategies that rely on the symmetry of our self-supervised loss in training and testing, in particular optimizing model parameters and/or the output of different tasks, thus leveraging their mutual interactions. The idea of jointly optimizing the system output, under all geometric and photometric constraints can be viewed as a dense generalization of classical bundle adjustment. We demonstrate the effectiveness of our method on KITTI and Cityscapes, where we outperform previous self-supervised approaches on multiple tasks. We also show good generalization for transfer learning in YouTube videos.
Code (0)
등록된 구현이 없습니다.
Tasks
Monocular Depth EstimationOptical Flow EstimationSelf-Supervised LearningTransfer LearningUnsupervised Monocular Depth EstimationSimilar Papers 제목 키워드 기반
$S^3$Net: Semantic-Aware Self-supervised Depth Estimation with Monocular Videos and Synthetic Data
Solving depth estimation with monocular cameras enables the possibility of widespread use of cameras as low-cost depth estimation sensors in applications such as autonomous driving and robotics. However, learning such a …
Autonomous DrivingDepth EstimationDomain AdaptationPanoptic SegmentationMGNet: Monocular Geometric Scene Understanding for Autonomous Driving
We introduce MGNet, a multi-task framework for monocular geometric scene understanding. We define monocular geometric scene understanding as the combination of two known tasks: Panoptic segmentation and self-supervised m…
Autonomous DrivingDepth EstimationGPUMonocular Depth Estimation+23D Consistency Optimization for Self-Supervised Monocular Video Depth Estimation
Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically treat video frames independently or rely o…
Multi-View 3D ReconstructionDepth EstimationSynthesizing Light Field Video from Monocular Video
The hardware challenges associated with light-field(LF) imaging has made it difficult for consumers to access its benefits like applications in post-capture focus and aperture control. Learning-based techniques which sol…
Self-Supervised LearningVideo ReconstructionS³Net: Semantic-Aware Self-supervised Depth Estimation with Monocular Videos and Synthetic Data
Solving depth estimation with monocular cameras enables the possibility of widespread use of cameras as low-cost depth estimation sensors in applications such as autonomous driving and robotics. In order to learn such a …
Autonomous DrivingDepth EstimationDomain Adaptation